OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show - Beritaja
OpenAI has offered a deeper look at its in-house AI inference hardware, Jalapeño, revealing early performance figures at the Hot Chips conference on Tuesday.
The company said tests conducted using SemiAnalysis’ InferenceX benchmark showed Jalapeño outperforming currently available leading inference hardware in two important areas: the number of tokens it can process for users and the amount of inference work it can deliver for each unit of electricity consumed.
Richard Ho, OpenAI’s head of hardware, described the results as a substantial improvement over existing systems during a media briefing. According to Ho, Jalapeño is designed to combine high processing efficiency with fast response times, allowing the system to handle large numbers of AI requests without sacrificing latency.
The comparison is particularly notable because OpenAI tested Jalapeño against a system based on Nvidia’s Blackwell architecture. However, the competitive landscape could look considerably different by the time OpenAI begins deploying its hardware at scale.
Ho said initial Jalapeño deployments are expected toward the end of 2026, although availability will initially be limited. Broader deployment is expected to take place during 2027.
OpenAI first revealed Jalapeño in October 2025. The chip is being developed in partnership with Broadcom, with OpenAI using its own AI models as part of the hardware development process. This development is part of a broader evolution in AI software and technology, where advances in computing infrastructure increasingly influence how AI systems are built and deployed.
That integrated strategy is particularly important for AI inference, where performance is affected by more than raw computing power. OpenAI says Jalapeño has been engineered to reduce delays during different stages of inference, including the initial prefill phase and the communication required as a model generates an answer.
The architecture also focuses on reducing unnecessary movement of data. OpenAI says model information, including the KV cache maintained during response generation, can remain close to the computing resources that need it. The system can then coordinate computing capacity, memory and networking according to the specific requirements of each inference stage.
The broader goal is to make AI inference both faster and more energy efficient, potentially allowing OpenAI to serve more users with the same amount of infrastructure and power.
If the early benchmark results translate into real-world deployments, Jalapeño could become an important part of OpenAI's strategy for controlling the infrastructure required to run increasingly demanding AI models.