OpenAI Releases First Benchmark Results for its Jalapeño AI Chip

OpenAI has published the first performance results for Jalapeño, its in-house AI inference processor designed to speed up model responses and lower server running costs. The benchmark data shows that the custom silicon delivers a major leap forward, yielding more AI work per watt of power.

Bar chart comparing peak per‑user decoding throughput across three model families: GPT-OSS 120B (green 1,459; blue 535), DeepSeek R1 670B (green 700; blue 169), Kimi K2.5 1T (green 694; blue 182); with 2.7x, 4.1x, and 3.8x annotations.

The launch marks a key shift for the ChatGPT creator. By designing its own hardware targeted specifically at running finished AI models rather than training them, OpenAI aims to lower operating expenses and handle heavy user traffic without relying solely on third-party graphics processors.

During public testing using the SemiAnalysis InferenceX benchmark suite, Jalapeño was pitted against leading commercial AI systems across several open-weight AI models, including GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. The results highlighted notable gains across the board:

  • Lower Latency: End-to-end response times dropped by 1.7 to 3.6 times compared to standard commercial setups.
  • Higher Efficiency: Jalapeño completed 1.5 to 1.9 times more total work per watt of electricity at peak capacity.
  • Interactive Gains: For rapid, real-time tasks like AI agent execution, performance jumped by 2.1 to 4.1 times.

The chip operates at a rated thermal design power of 700 watts, though real-world testing showed it sustained power consumption under 550 watts during standard workloads. On the 1-trillion-parameter Kimi K2.5 model, the system returned responses 3.4 times faster than rival configurations while using significantly less power.

Building custom silicon usually takes years of manual engineering. To speed up development, OpenAI used its own AI models to help design and test Jalapeño, taking the processor from concept to tapeout in just nine months.

OpenAI plans to begin deploying Jalapeño processors directly into its production data centers before the end of the year. The initial rollout will handle real-world traffic alongside existing hardware from external partners like Nvidia.

Want to see more of our stories on Google?

Add iPhone in Canada as a Preferred Source on Google

P.S. Want to keep this site truly independent? Support us by buying us a beer, treating us to a coffee, or shopping through Amazon here. Links in this post are affiliate links, so we earn a tiny commission at no charge to you. Thanks for supporting independent Canadian media!

Subscribe
Notify of
guest
0 Comments
Oldest
Newest Most Voted
0
Would love your thoughts, please comment.x
()
x