OpenAI says Jalapeño beat Nvidia GB300 in tests, but benchmark limits remain
OpenAI says its in-house Jalapeño inference chip, developed with Broadcom, outperformed Nvidia’s GB300 on power efficiency and response time in internal testing, according to remarks by chip lead Richard Ho at Stanford University’s Hot Chips conference. The chip is designed for inference rather than model training and was formally introduced in June after a development cycle that OpenAI says took nine months from project start to tape-out. A 128-chip Jalapeño system delivers 1.7 exaFLOPS of 4-bit MXFP4 compute, or about 13.4 petaFLOPS per chip, and includes 27.5 TB of HBM4 memory. Technical analysis cited by The Register said the system posted 1.5x to 1.9x higher peak throughput and 1.7x to 3.6x lower end-to-end latency in OpenAI’s custom InferenceX benchmark, with even larger gains in ultra-low-latency scenarios. Bloomberg, however, said the results come with at least four caveats: Jalapeño was not compared with Nvidia’s newer Vera Rubin chips; speculative decoding was excluded from the benchmark; the system trails Nvidia GB200 NVL72 and GB300 NVL72 racks in raw compute and memory capacity; and OpenAI has not disclosed actual power draw. The report added that volumes will remain limited in 2026, with mass production only expected in 2027.








