News

OpenAI’s Jalapeño – update 

More details from Hot Chips—OpenAI challenges Nvidia.

Jon Peddie

OpenAI built a chip. Not a demo, not a paper design, a working ASIC called Jalapeño, taped-out in November 2025 and already running real inference workloads today. OpenAI published its benchmark numbers, which indicate Jalapeño delivers more tokens per watt than anything else currently shipping, built by a company that had never designed a chip before. We revisit its architecture, the numbers, and what a credible first-generation ASIC means for the merchant silicon business.

OpenAI partnered with Broadcom in June 2026 to reveal Jalapeño, an inference-only ASIC built from a blank sheet starting mid-2024. Tape-out landed roughly 16 months later, an unusually fast cycle for a first chip. First-generation accelerators typically underperform. Jalapeño doesn’t follow that pattern: SemiAnalysis benchmarked it against Nvidia, AMD, and Google chips across multiple open models and found it leading on throughput per watt in nearly every scenario tested.

The chip isn’t specialized for OpenAI’s own models, a common misreading of the announcement. It runs DeepSeek R1, Kimi K2.5, and GPT-OSS, with strong results across all three, plus a Codex-ported build of Doom as a demonstration. On DeepSeek R1 at low concurrency, Jalapeño hits over 700 tokens per second per user, achieved with single-token prediction alone—no speculative decoding, no prefill-decode disaggregation, techniques competing chips lean on for their best numbers.

Power drives every design decision here. OpenAI runs power-constrained today, not budget or floor-space-constrained, making tokens per megawatt the metric that matters. Nvidia’s own Hot Chips 2026 talk made the same point: “The data center is power-limited today.” Grid interconnection now moves slower than hardware deployment, pushing operators toward behind-the-meter generation, gas turbines, and on-site power sitting outside utility control entirely.

Throughput-per-kW Frontier

Figure 1. InferenceX GPT‑OSS‑120B nominal 8K/1K STP package TDP: Jalapeño 700 W; GB200 1,200 W. (Source: OpenAI)

On that efficiency metric, Jalapeño beats Vera Rubin’s published July 2026 numbers, Nvidia’s own next-generation chip, despite Rubin taping-out a month earlier and reaching customers first. OpenAI’s own published data adds independent confirmation: 1.5 to 1.9 times more work per watt at peak throughput across GPT-OSS, DeepSeek R1, and Kimi K2.5, with the advantage widening further on OpenAI’s internal frontier models. On total cost of ownership, Jalapeño and Rubin run close to even, though Jalapeño reaches that parity without speculative decoding, meaning real headroom remains once that optimization lands.

Real caveats apply. The models benchmarked so far skip the largest frontier releases like DeepSeek V4 Pro. And Jalapeño remains an engineering sample. Production ramps through 2027, with OpenAI targeting 100 MW of deployed capacity as the near-term goal, and hardware supply and data center operations standing as the real constraints ahead, not the chip design itself.

The architecture explains the results. Jalapeño skips prefill-decode disaggregation entirely, keeping one unified chip pool instead of splitting hardware by workload phase. Production traffic shifts constantly between prefill-heavy and decode-heavy demand, and a fixed split strands capacity whenever that ratio moves. A unified pool trades some theoretical peak efficiency for consistently higher real-world utilization. Out-of-order cores with local cache replace the software-managed scratchpads standard on GPUs and TPUs, cutting the fixed latencies that force other architectures to hide overhead behind larger batch sizes.

AI wrote much of the software stack. OpenAI’s Codex generated kernels for models the chip’s original production plan never included, reaching high performance within two months—no kernel engineers required for that work. Select attention and mixture-of-experts blocks ran 1.5 to 1.8 times faster in AI-generated form than the existing human-written versions. The chip’s design process ran the same way: Earlier OpenAI models helped bring up the silicon itself, creating a real feedback loop between the models a company builds and the chips it designs to run them.

Two chip generations already sit in development behind Jalapeño. Gen 2 is well underway; Gen 3 is taking shape. OpenAI continues buying Nvidia and other merchant silicon for both training and inference alongside its own chip, not replacing external suppliers, but adding a credible internal option that raises the bar competitors now have to clear.

What do we think?

A credible first-generation ASIC from a company with zero chip design history is a genuine surprise, and the power-efficiency numbers hold up under independent testing, not just OpenAI’s own claims. The real question isn’t whether Jalapeño works; it clearly does. It’s whether OpenAI can scale manufacturing and data center deployment fast enough to matter before Rubin’s software matures to match it.

Inflection point. A frontier AI lab designing chips that beat merchant silicon on power efficiency marks a real inflection point for the semiconductor industry. If Jalapeño’s trajectory holds through Gen 2 and Gen 3, the AI labs building the largest models stop being Nvidia’s customers and start being its competitors, at least for their own workloads. That shift changes who captures the margin in AI infrastructure, and it started with a company that had never taped-out a chip before this one.

LIKE WHAT YOU’RE READING? TELL YOUR FRIENDS; WE DO THIS EVERY DAY, ALL DAY.