OpenAI Jalapeño: OpenAI's First Self-Developed Inference Chip Explained
Quick verdict
OpenAI's first self-developed inference chip Jalapeño, co-built with Broadcom on TSMC 3nm, shows 1.5–1.9x per-watt throughput vs NVIDIA GB200/GB300 and roughly 1.7–3.6x lower end-to-end latency in SemiAnalysis InferenceX tests. Small-batch deployment planned for late 2026, larger scale in 2027.
OpenAI has shown the first public benchmarks for Jalapeño, its first self-developed AI inference chip. Co-built with Broadcom and manufactured on TSMC's 3nm process, Jalapeño is targeted at one thing: serving AI models cheaply and quickly at scale.
It's a significant step in OpenAI's effort to reduce dependence on NVIDIA and control its own inference economics — and it overlaps directly with the same price/scaling pressure that just hit DeepSeek V4 and GPT-5.6 Sol on the API side.
What Jalapeño Is
- Type: Custom AI inference ASIC (specialized chip, not a GPU)
- Partner: Broadcom (co-developed)
- Manufacturing: TSMC 3nm process
- Design-to-tapeout: about 9 months (industry-record speed)
- Power rating: 700W, sustained power ≤550W during test workloads
- Deployment plan: small batch in late 2026, larger scale in 2027
The Benchmarks
OpenAI tested Jalapeño on SemiAnalysis's InferenceX benchmark (a public, vendor-neutral tool for LLM serving), comparing it to NVIDIA's GB200 and GB300 systems across three open models:
| Model | Compared to | Per-watt throughput | End-to-end latency |
|---|---|---|---|
| GPT-OSS 120B | GB200 | ~1.9× higher | ~28–35% (≈3–3.6× lower) |
| DeepSeek R1 670B | GB300 | ~1.7× higher | roughly 59% (≈1.7× lower) |
| Kimi K2.5 1T | GB300 | ~1.5× higher | between GB200 and GB300 range |
⚠️ Caveats: the benchmarks and methodology were chosen by OpenAI. There is no independent third-party verification yet, and OpenAI hasn't published pricing or total cost of ownership.
Why It Matters
- Lower inference cost: per-watt gains of 1.5–1.9× translate directly to cheaper serving for models like GPT-5.6 Sol and open-weight models OpenAI also serves.
- Lower latency: roughly half (or better) of GB300 end-to-end latency means faster responses for long-context and agent workloads.
- NVIDIA optionality: OpenAI now has a credible second source for inference compute, which puts pressure on NVIDIA pricing and supply.
What's Still Unclear
- No third-party verification of the InferenceX numbers
- No chip price or TCO figures — "1.7× per-watt" doesn't automatically mean cheaper tokens at scale
- Scale timing: real volume isn't until 2027, so this doesn't change NVIDIA procurement decisions today
- Software stack: how well Jalapeño integrates with the wider OpenAI / Azure stack is still being validated
Bottom Line
Jalapeño is a credible first move — OpenAI's inference cost structure and latency ceiling get a real upgrade path. But until third-party benchmarks and production-scale numbers land, treat this as an early signal, not a market shift.
If you're an AI cloud buyer, don't change your procurement yet. If you're building on OpenAI APIs, watch for latency and price changes that could trickle through in 2027.
Related Articles
Related Articles
NVIDIA Acquires Hugging Face for $12.93B: What It Means for Open AI Models
NVIDIA agreed to buy Hugging Face for $12,930,300,000 on September 3, 2026. What the deal covers, what NVIDIA promised about openness, and why it matters for developers.
AI Solves the Navier-Stokes Millennium Prize Problem: What OpenAI's Agents Actually Did
OpenAI's internal agent fleet produced a Navier-Stokes blow-up solution in ~88 hours using 2.7M agent messages and 130B output tokens, with Lean verification by GPT-6 Astra. Full breakdown.
OpenAI Cuts GPT-5.6 Sol API Price by 20%+: What Developers Should Know
OpenAI lowered GPT-5.6 Sol developer API pricing by more than 20% starting August 21, 2026. See the approximate new price, why it matters, and how it changes the AI pricing landscape.
GPT Image 2.5 Review: OpenAI's Faster, Sharper Image Model (2026)
GPT Image 2.5 reviewed — 50% lower latency, better multi-edit consistency, the new Sketch feature, templates, and the Flare vs Sunburst API variants. Facts verified from OpenAI.