OpenAI’s Jalapeño Chip Beats Nvidia’s GB300 on AI Speed and Efficiency

OpenAI Jalapeño chip

OpenAI has released its first detailed performance results for Jalapeño, its custom AI inference chip developed with Broadcom, and the early numbers show a significant advantage over Nvidia’s existing systems in key inference workloads.

OpenAI says Jalapeño delivered between 1.5 and 1.9 times more AI work per watt than the comparison systems, while reducing end-to-end latency by 1.7 to 3.6 times. The results come from tests involving GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 1T.

The results are significant because OpenAI is building Jalapeño specifically around the economics of serving AI models at scale. Unlike a general-purpose accelerator, the chip is designed for inference, where models generate responses for users after they have been trained.

Jalapeño focuses on the economics of AI inference

For OpenAI, inference is becoming an increasingly important part of its computing requirements as ChatGPT, Codex and its API serve growing numbers of users.

Jalapeño is designed to address two competing requirements at once: serving large numbers of requests efficiently while keeping response times low.

OpenAI says the chip achieved both in its latest testing. Across the three models, it reported between 1.5 and 1.9 times more AI work per watt at peak throughput and between 1.7 and 3.6 times lower end-to-end latency than the systems used for comparison.

For interactive AI applications, the latency improvement can be particularly important. OpenAI says Jalapeño delivered between 2.1 and 4.1 times higher performance for highly interactive workloads.

That combination could help OpenAI serve more AI requests while using less power for the same amount of useful work.

The 700-watt figure needs some context

Jalapeño has a 700-watt power rating, but that does not mean the chip continuously consumed 700 watts throughout OpenAI’s tests.

According to the company’s measurements, sustained power remained at or below 550 watts on the workloads it tested. The distinction is important when comparing the chip’s efficiency because the 700-watt figure represents its rated power rather than its measured sustained consumption.

Power efficiency matters enormously in large AI data centers. Every additional watt required to serve an inference workload adds to the cost of operating the infrastructure, particularly when thousands of accelerators are running continuously.

That gives custom inference silicon a potentially important role for companies operating AI services at enormous scale.

Jalapeño was tested across several large models

OpenAI did not limit its evaluation to its own models.

The company tested Jalapeño with GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, giving the chip a range of workloads spanning models developed both inside and outside OpenAI.

OpenAI says Jalapeño remained on the performance-efficiency frontier across different operating points for all three models.

That is notable because custom accelerators often gain an advantage by being tightly optimized around a particular workload. OpenAI is instead presenting Jalapeño as a platform capable of handling different large language models.

The company says the architecture was designed around the memory movement, networking and serving patterns that matter most for modern LLM inference.

This is not a Vera Rubin comparison

There is an important limitation to the results.

OpenAI’s latest comparison does not establish that Jalapeño is faster or more efficient than Nvidia’s newest Vera Rubin platform. Vera Rubin represents Nvidia’s newer generation of AI infrastructure, while the reported Jalapeño tests compare against existing Nvidia systems.

That means the results should not be interpreted as OpenAI overtaking Nvidia across the entire AI hardware market.

Instead, they show that a purpose-built inference accelerator can compete strongly with established GPU-based systems when it is optimized specifically for the workloads an AI company needs to serve.

OpenAI’s own testing is also only one part of the picture. The company selected the workloads and comparison systems, so independent testing and large-scale deployment will provide a more complete assessment of Jalapeño’s real-world advantage.

Jalapeño is built for inference, not training

Another important distinction is what the chip is designed to do.

Jalapeño is an application-specific processor built for large language model inference. It is not intended to replace the general-purpose accelerators OpenAI uses to train its models.

That makes the chip complementary to OpenAI’s existing compute infrastructure rather than a complete replacement for Nvidia hardware.

Training frontier models requires enormous amounts of flexible compute, while inference involves repeatedly serving already-trained models. A company with OpenAI’s scale can therefore benefit from building specialized hardware for the second stage without abandoning the GPUs needed for training.

OpenAI is still expected to rely heavily on Nvidia hardware alongside its own silicon.

OpenAI is already planning the next generations

Jalapeño is also not being presented as a one-off experiment.

OpenAI describes it as the beginning of a multi-generation inference platform. The company says it plans to ramp the chip into its infrastructure and continue developing successors as its AI workloads evolve.

The first deployments are expected to begin on a limited basis toward the end of 2026, with broader deployment planned for 2027.

That longer roadmap could ultimately matter more than the first-generation benchmark numbers. If OpenAI can repeatedly design specialized hardware around the behavior of its models and serving systems, it could gradually reduce the amount of inference capacity it needs to purchase from outside chip suppliers.

Nvidia is not the only company facing the custom-chip push

OpenAI’s move is part of a much broader shift across the AI industry.

Major AI companies and cloud providers are increasingly developing custom accelerators to reduce costs and gain more control over their computing infrastructure.

Google has its TPU family, Amazon has Trainium and Inferentia, while Microsoft and Meta have also developed custom AI silicon. The goal is similar across the industry: optimize hardware around specific AI workloads instead of relying exclusively on general-purpose accelerators.

For OpenAI, Jalapeño adds another layer to that strategy.

The company already controls its models, products and software stack. Building its own inference hardware gives it greater control over the infrastructure underneath those services as well.

Jalapeño therefore does not spell the end of Nvidia’s role at OpenAI. But the early results show why the world’s largest AI companies are increasingly willing to invest billions in custom silicon: at inference scale, even modest improvements in performance per watt can become a major economic advantage.

Leave a Reply

Your email address will not be published. Required fields are marked *

Previous Post
Apple M6 and M5 Ultra

Apple Unveils M6 and M5 Ultra With Huge AI Performance Gains

Related Posts