OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

Published: 2026-08-26

OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
At the Hot Chips conference on Tuesday, OpenAI shared a more detailed look at Jalapeño, including the first batch of benchmark results for the new system. Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the-art inference processors. “The bottom line is that the results show a very, very significant performance advance over state of the art,” said Richard Ho, OpenAI’s head of hardware, in a press call. “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It’s very efficient to serve a lot of customers, but it can also be very low latency.”

Originally sourced from Techcrunch

Read the full story on Global Insight Daily