OpenAI says Jalapeño beat Nvidia GB300 in tests, but benchmark limits remain

OpenAI says Jalapeño beat Nvidia GB300 in tests, but benchmark limits remain

N
News Editor
2026-08-26 04:52:38
OpenAI says its in-house Jalapeño inference chip, developed with Broadcom, outperformed Nvidia’s GB300 on power efficiency and response time in internal testing, according to remarks by chip lead Richard Ho at Stanford University’s Hot Chips conference. The chip is designed for inference rather than model training and was formally introduced in June after a development cycle that OpenAI says took nine months from project start to tape-out. A 128-chip Jalapeño system delivers 1.7 exaFLOPS of 4-bit MXFP4 compute, or about 13.4 petaFLOPS per chip, and includes 27.5 TB of HBM4 memory. Technical analysis cited by The Register said the system posted 1.5x to 1.9x higher peak throughput and 1.7x to 3.6x lower end-to-end latency in OpenAI’s custom InferenceX benchmark, with even larger gains in ultra-low-latency scenarios. Bloomberg, however, said the results come with at least four caveats: Jalapeño was not compared with Nvidia’s newer Vera Rubin chips; speculative decoding was excluded from the benchmark; the system trails Nvidia GB200 NVL72 and GB300 NVL72 racks in raw compute and memory capacity; and OpenAI has not disclosed actual power draw. The report added that volumes will remain limited in 2026, with mass production only expected in 2027.

OpenAI’s Jalapeño inference chip, built with Broadcom, outperformed Nvidia’s current flagship GB300 on power efficiency and response time in testing, according to Richard Ho, OpenAI’s chip lead, speaking at Stanford University’s Hot Chips conference.

Bloomberg said the result comes with at least four caveats that were not fully laid out alongside the performance claims. The advantage shown so far is centered on performance per watt and latency, while large-scale deployment is still some distance away.

Built for inference, not training

Jalapeño is an inference-only chip developed jointly by OpenAI and Broadcom. It is meant to handle question answering and task execution, rather than training models from scratch, which requires a different design approach.

The partnership had already been disclosed last year, and the chip was formally introduced in June after what OpenAI described as an unusually fast development cycle. In testing, OpenAI ran Jalapeño on its own small open-source model as well as third-party models from DeepSeek and Moonshot AI. Ho said the biggest advantage showed up on Moonshot’s Kimi model.

Ho also said internal testing showed strong performance on a larger advanced model that OpenAI has not yet released.

128-chip system delivers 1.7 exaFLOPS of MXFP4 compute

On specifications, a system made up of 128 Jalapeño chips provides 1.7 exaFLOPS of 4-bit MXFP4 compute, which works out to roughly 13.4 petaFLOPS per chip.

The Register, in a technical analysis, said that in OpenAI’s custom InferenceX benchmark, the system produced 1.5x to 1.9x higher peak throughput than the comparison rack system, with end-to-end latency reduced by 1.7x to 3.6x. In ultra-low-latency scenarios, it was 2.1x to 4.1x faster. The tested models included GPT-OSS-120B, DeepSeek R1, and Kimi K2.5.

Each system includes 27.5 TB of HBM4 high-bandwidth memory, averaging 216 GB per chip. The report said total memory bandwidth is close to 2 petabytes per second. It also includes a large SRAM cache for KV cache storage and intermediate compute results.

Nine months from project start to tape-out

From project launch to tape-out, Jalapeño took nine months, with OpenAI saying AI models helped speed up the design process. The report described tape-out as the final step before a chip design is fixed and sent to a foundry for production.

The second-generation Jalapeño has already moved into the back-end stage, and third-generation concepts are now in design. OpenAI plans to deploy Jalapeño later in 2026 and decide at that point which models will run on it. The company said the goal is to cut data center electricity costs.

Bloomberg’s four caveats

Bloomberg said the headline result needs to be read with four important limits in mind.

First, Jalapeño was not tested against Nvidia’s newer Vera Rubin chips, which have only just begun shipping. The comparison was made against the older GB300 generation.

Second, the benchmark excluded speculative decoding. The report described that as a method where a smaller model drafts outputs and a larger model verifies them, a technique widely used in the industry to raise inference speed. Leaving it out makes the comparison relatively more favorable to Jalapeño’s architecture.

Third, against Nvidia’s GB200 NVL72 and GB300 NVL72 racks, Jalapeño has 1.46x to 2x less raw compute and about 10% less memory capacity, while leading by nearly 20% only in memory bandwidth. In other words, its edge is in efficiency per watt and latency rather than total compute scale.

Fourth, OpenAI has not disclosed Jalapeño’s actual power consumption. External estimates cited in the report put power savings at 40% to 60% versus a typical GPU. The report also said the chip will not enter mass production until 2027, with only limited supply in 2026, leaving Nvidia’s near-term data center shipment position largely intact.

By Bloomberg’s framing, the result looks more like a strong lab scorecard than a sign of immediate large-scale rollout.

OpenAI says it still needs Nvidia

Ho did not present Jalapeño as a replacement for Nvidia across the board. He said Nvidia remains a 「very good partner」 and that OpenAI will still need 「a lot of Nvidia」.

At the same time, OpenAI continues to use technology from Cerebras Systems to run some models. Ho said, 「Our demand for compute is so large,」 and added that the company has already signed with multiple suppliers.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
30

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.