OpenAI Jalapeño: OpenAI's First Self-Developed Inference Chip Explained
Quick verdict
OpenAI's first self-developed inference chip Jalapeño, co-built with Broadcom on TSMC 3nm, shows 1.5–1.9x per-watt throughput vs NVIDIA GB200/GB300 and roughly 1.7–3.6x lower end-to-end latency in SemiAnalysis InferenceX tests. Small-batch deployment planned for late 2026, larger scale in 2027.
OpenAI has shown the first public benchmarks for Jalapeño, its first self-developed AI inference chip. Co-built with Broadcom and manufactured on TSMC's 3nm process, Jalapeño is targeted at one thing: serving AI models cheaply and quickly at scale.
It's a significant step in OpenAI's effort to reduce dependence on NVIDIA and control its own inference economics — and it overlaps directly with the same price/scaling pressure that just hit DeepSeek V4 and GPT-5.6 Sol on the API side.
What Jalapeño Is
- Type: Custom AI inference ASIC (specialized chip, not a GPU)
- Partner: Broadcom (co-developed)
- Manufacturing: TSMC 3nm process
- Design-to-tapeout: about 9 months (industry-record speed)
- Power rating: 700W, sustained power ≤550W during test workloads
- Deployment plan: small batch in late 2026, larger scale in 2027
The Benchmarks
OpenAI tested Jalapeño on SemiAnalysis's InferenceX benchmark (a public, vendor-neutral tool for LLM serving), comparing it to NVIDIA's GB200 and GB300 systems across three open models:
| Model | Compared to | Per-watt throughput | End-to-end latency |
|---|---|---|---|
| GPT-OSS 120B | GB200 | ~1.9× higher | ~28–35% (≈3–3.6× lower) |
| DeepSeek R1 670B | GB300 | ~1.7× higher | roughly 59% (≈1.7× lower) |
| Kimi K2.5 1T | GB300 | ~1.5× higher | between GB200 and GB300 range |
⚠️ Caveats: the benchmarks and methodology were chosen by OpenAI. There is no independent third-party verification yet, and OpenAI hasn't published pricing or total cost of ownership.
Why It Matters
- Lower inference cost: per-watt gains of 1.5–1.9× translate directly to cheaper serving for models like GPT-5.6 Sol and open-weight models OpenAI also serves.
- Lower latency: roughly half (or better) of GB300 end-to-end latency means faster responses for long-context and agent workloads.
- NVIDIA optionality: OpenAI now has a credible second source for inference compute, which puts pressure on NVIDIA pricing and supply.
What's Still Unclear
- No third-party verification of the InferenceX numbers
- No chip price or TCO figures — "1.7× per-watt" doesn't automatically mean cheaper tokens at scale
- Scale timing: real volume isn't until 2027, so this doesn't change NVIDIA procurement decisions today
- Software stack: how well Jalapeño integrates with the wider OpenAI / Azure stack is still being validated
Bottom Line
Jalapeño is a credible first move — OpenAI's inference cost structure and latency ceiling get a real upgrade path. But until third-party benchmarks and production-scale numbers land, treat this as an early signal, not a market shift.
If you're an AI cloud buyer, don't change your procurement yet. If you're building on OpenAI APIs, watch for latency and price changes that could trickle through in 2027.
Related Articles
Related Articles
OpenAI Cuts GPT-5.6 Sol API Price by 20%+: What Developers Should Know
OpenAI lowered GPT-5.6 Sol developer API pricing by more than 20% starting August 21, 2026. See the approximate new price, why it matters, and how it changes the AI pricing landscape.
GPT-5.6 Sol, Terra & Luna: Complete Guide to OpenAI's New Model Family (2026)
Everything about GPT-5.6 Sol, Terra, and Luna — specs, pricing (Luna down 80%), performance benchmarks, new API features, and how they compare to Claude Fable 5 and DeepSeek V4.
Sora: OpenAI's AI Video Generator — Complete Review 2026
A comprehensive review of OpenAI Sora's legacy. Sora's web and app shut down in April 2026 with the API ending September 2026. Explore its impact, quality, and the best alternatives for AI video generation.
AI for Customer Service: Best Tools and Strategies in 2026
Discover how AI is transforming customer service in 2026. Compare the best AI customer service tools, chatbots, and platforms for businesses of all sizes.