
The news of August 26thOpenAI Publication of the down payment self-study reasoningchip Galapeño ' s first results were measured across the three indicators of throughput, energy efficiency and delayNvidia GB200 with GB300. According to the results of the OpenAI and SemiAnalysis tests on the InferenceX baseline:
On GPT-OSS 120B, single-user velocity of 1459 tokens/s is 2.7 times greater than GB200 535 tokens/s
On DeepSeek R1 670B, the end-to-end delay of 1.65 seconds compared to 5.99 seconds of GB300; based on GB300 ' s fastest decoder speed, 104.3 times more per kilowatt
(a) On Kimi K2.5 1T, the peak is about 1.5 times higher per watt-off and the delay is about 3.4 times lower
THE TOTAL WORKLOAD OF THE THREE MODELS IS 1.5 TO 1.9 TIMES HIGHER AND THE PERFORMANCE OF HIGH-INTERACTIVE SCENES IS 2.1 TO 4.1 TIMES HIGHER。
Galapeño ' s nominal power consumption of 700W, which is only half of GB300; SemiAnalysis estimates the total cost of ownership per chip at approximately $1.56 per hour, which is essentially the same as the British Virgin H100, which is less than the $3.61 of Vera Rubin。
OpenAI states that Jalapeño only took nine months from the design to the current film, and in the process using the self-study model GPT-Astra and Codex accelerated circuit design and validation, AI supported the reduction of 8% for the SIMD unit and 10% for the matrix engine; on the partial attention and MoE module, AI produced a kernel 1.5 to 1.8 times faster than human experts。
Last October, OpenAI signed a 10-GW Custom Accelerator Cooperation Agreement with Chase. Jalapeño was the first generation product of the road map, Gen 2 has been developed in depth and Gen 3 is taking shape. Officially, Galapeño will be deployed into its own computing infrastructure by the end of this year。