OpenAI's Jalapeño chip is built for fast inference at scale, benchmarks show

2026-08-26·8 min read

On August 25, 2026, at the highly anticipated Hot Chips conference, OpenAI publicly revealed the detailed architecture of Jalapeño, its custom inference chip, and published its first batch of benchmark results. Tested on SemiAnalysis' InferenceX benchmark, Jalapeño beat currently available state-of-the-art inference processors on two key metrics — tokens served per user and throughput per kilowatt of power — and the comparison target was Nvidia's Blackwell system. OpenAI's head of hardware, Richard Ho, said plainly in a press call: 'The bottom line is that the results show a very, very significant performance advance over state of the art.' For a company that has historically run its models mostly on Nvidia GPUs, this is the most powerful public proof yet of its 'de-Nvidia-fication' strategy.

First, understand Jalapeño's positioning. It is not a training chip for building large models — it is designed specifically for inference, the computing that actually runs the model and generates answers in the background when a user asks ChatGPT a question. As ChatGPT's user base has grown exponentially, inference cost has become one of OpenAI's largest expense items. Ho stressed in the press call that Jalapeño has two design goals: serving more AI workloads per unit of power, and returning responses with very low latency. In other words, OpenAI wants this chip to support high-concurrency access from massive numbers of users while keeping each user's wait time short. In the InferenceX benchmark, Jalapeño outperformed Blackwell on both dimensions. Ho acknowledged, however, that the comparison has a time gap — Nvidia's Blackwell is currently shipping hardware, while Jalapeño will not deploy until late 2026 'in very small volumes', with significant deployment coming in 2027. By then, Nvidia's next-generation products may also be in place, and the competitive landscape may not look the same as today.

The R&D backdrop of Jalapeño is equally noteworthy. The chip was first announced in October 2025, developed by OpenAI in close collaboration with Broadcom, and OpenAI's own models assisted in the chip's development — using AI to design AI hardware is itself an expression of full-stack innovation. OpenAI plans to make Jalapeño a multigenerational platform, allowing AI products, models, chips, and memory to be developed in concert. Because of this full-stack approach, OpenAI was able to target the phases of inference that often create bottlenecks: Jalapeño is specifically designed to minimize delays during the prefill phase — the most time-consuming part when a model processes a prompt — and the communication phase, the overhead of data exchange when multiple chips work together. Ho said OpenAI hopes this software-hardware co-design will let Jalapeño serve customers at scale efficiently while keeping latency low, significantly reducing unit compute cost without sacrificing experience.

From an industry-structure perspective, Jalapeño's debut is a landmark in the wave of AI compute autonomy. Over the past two years, nearly every leading AI company has been building custom silicon to escape dependence on a single chip supplier and cut soaring compute bills: Google has TPU, Amazon has Trainium and Inferentia, Meta is advancing its own chip project, and Microsoft is developing custom hardware with OpenAI. As the company most heavily dependent on Nvidia, OpenAI's chip progress has long been seen as a bellwether. Publishing these benchmarks sends several key signals: first, custom inference chips can now compete head-on with the traditional champion on efficiency; second, software-hardware co-design is becoming the new high ground of AI infrastructure competition; third, energy-efficiency metrics like throughput per kilowatt are replacing raw compute numbers as the new yardstick for AI chips — because in an era of soaring electricity costs, efficiency is cost, and cost is competitiveness. Of course, there is a long road from benchmarks to large-scale stable deployment, and whether Jalapeño can truly carry ChatGPT's inference load in 2027 remains to be seen. But either way, the era of 'Nvidia dominance' is being rewritten, chip by chip, by these challengers.

📌 Source: TechCrunch (August 25, 2026) — 'OpenAI's Jalapeño chip is built for fast inference at scale, benchmarks show' by Russell Brandom. Link: techcrunch.com/2026/08/25/openais-jalapeno-chip-is-built-for-fast-inference-at-scale-benchmarks-show/

🤔 Frequently Asked Questions

Q1: What is Jalapeño?

Jalapeño is OpenAI's custom inference chip developed with Broadcom, designed for inference computing — running AI models — rather than training. First announced in October 2025, it aims to support massive inference loads for products like ChatGPT while lowering unit compute cost.

Q2: How did the benchmarks go?

On SemiAnalysis' InferenceX benchmark, Jalapeño beat the current state-of-the-art inference processors — compared against Nvidia Blackwell — on tokens per user and throughput per kilowatt. OpenAI's hardware chief called it 'a very, very significant performance advance over state of the art.'

Q3: When will the chip deploy?

According to OpenAI hardware chief Richard Ho, Jalapeño will deploy in very small volumes at the end of 2026, with more significant deployment coming in 2027.

Q4: Why is OpenAI building its own chip?

The core reasons are cost and autonomy: as ChatGPT's user base exploded, inference cost became one of the largest expenses; custom silicon also reduces dependence on a single supplier and enables software-hardware co-design that targets inference bottlenecks, creating an efficiency advantage.

🛠️ Recommended Tools

  • Text Summarizer - Quickly distill core findings from chip benchmark reports, technical whitepapers and industry analyses
  • API Tester - Test the response speed and stability of various AI inference APIs and experience inference performance differences yourself
  • PDF Summarizer - One-click extraction of key points from PDF reports and papers from conferences like Hot Chips to stay on top of chip frontier developments

Summary

OpenAI's Jalapeño benchmark data revealed at Hot Chips is an important footnote in the AI chip competitive landscape. It proves that custom inference chips can already compete head-on with Nvidia Blackwell on efficiency and throughput, and demonstrates the profound impact of software-hardware co-design on AI infrastructure. For OpenAI, Jalapeño means control over inference costs and security in the compute supply chain; for the industry, 'throughput per kilowatt' is becoming the new yardstick of competition, signaling that the AI race is extending from model capability to compute efficiency. Of course, from paper benchmarks to large-scale deployment in 2027, countless engineering challenges remain to be proven. But one clear trend is irreversible: AI companies are no longer content to rent someone else's compute — owning the chip is becoming a key part of owning AI's future.