OpenAI’s Jalapeño inference chip, built with Broadcom, outperformed Nvidia’s current flagship GB300 on power efficiency and response time in testing, according to Richard Ho, OpenAI’s chip lead, speaking at Stanford University’s Hot Chips conference.
Bloomberg said the result comes with at least four caveats that were not fully laid out alongside the performance claims. The advantage shown so far is centered on performance per watt and latency, while large-scale deployment is still some distance away.
Built for inference, not training
Jalapeño is an inference-only chip developed jointly by OpenAI and Broadcom. It is meant to handle question answering and task execution, rather than training models from scratch, which requires a different design approach.
The partnership had already been disclosed last year, and the chip was formally introduced in June after what OpenAI described as an unusually fast development cycle. In testing, OpenAI ran Jalapeño on its own small open-source model as well as third-party models from DeepSeek and Moonshot AI. Ho said the biggest advantage showed up on Moonshot’s Kimi model.
Ho also said internal testing showed strong performance on a larger advanced model that OpenAI has not yet released.
128-chip system delivers 1.7 exaFLOPS of MXFP4 compute
On specifications, a system made up of 128 Jalapeño chips provides 1.7 exaFLOPS of 4-bit MXFP4 compute, which works out to roughly 13.4 petaFLOPS per chip.
The Register, in a technical analysis, said that in OpenAI’s custom InferenceX benchmark, the system produced 1.5x to 1.9x higher peak throughput than the comparison rack system, with end-to-end latency reduced by 1.7x to 3.6x. In ultra-low-latency scenarios, it was 2.1x to 4.1x faster. The tested models included GPT-OSS-120B, DeepSeek R1, and Kimi K2.5.
Each system includes 27.5 TB of HBM4 high-bandwidth memory, averaging 216 GB per chip. The report said total memory bandwidth is close to 2 petabytes per second. It also includes a large SRAM cache for KV cache storage and intermediate compute results.
Nine months from project start to tape-out
From project launch to tape-out, Jalapeño took nine months, with OpenAI saying AI models helped speed up the design process. The report described tape-out as the final step before a chip design is fixed and sent to a foundry for production.
The second-generation Jalapeño has already moved into the back-end stage, and third-generation concepts are now in design. OpenAI plans to deploy Jalapeño later in 2026 and decide at that point which models will run on it. The company said the goal is to cut data center electricity costs.
Bloomberg’s four caveats
Bloomberg said the headline result needs to be read with four important limits in mind.
First, Jalapeño was not tested against Nvidia’s newer Vera Rubin chips, which have only just begun shipping. The comparison was made against the older GB300 generation.
Second, the benchmark excluded speculative decoding. The report described that as a method where a smaller model drafts outputs and a larger model verifies them, a technique widely used in the industry to raise inference speed. Leaving it out makes the comparison relatively more favorable to Jalapeño’s architecture.
Third, against Nvidia’s GB200 NVL72 and GB300 NVL72 racks, Jalapeño has 1.46x to 2x less raw compute and about 10% less memory capacity, while leading by nearly 20% only in memory bandwidth. In other words, its edge is in efficiency per watt and latency rather than total compute scale.
Fourth, OpenAI has not disclosed Jalapeño’s actual power consumption. External estimates cited in the report put power savings at 40% to 60% versus a typical GPU. The report also said the chip will not enter mass production until 2027, with only limited supply in 2026, leaving Nvidia’s near-term data center shipment position largely intact.
By Bloomberg’s framing, the result looks more like a strong lab scorecard than a sign of immediate large-scale rollout.
OpenAI says it still needs Nvidia
Ho did not present Jalapeño as a replacement for Nvidia across the board. He said Nvidia remains a 「very good partner」 and that OpenAI will still need 「a lot of Nvidia」.
At the same time, OpenAI continues to use technology from Cerebras Systems to run some models. Ho said, 「Our demand for compute is so large,」 and added that the company has already signed with multiple suppliers.

