Performance Review: OpenAI’s Jalapeño Chip In Action
AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: Performance Review: OpenAI’s Jalapeño Chip In Action on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

OpenAI released early performance data for its Jalapeño inference chip, showing notable improvements in efficiency and latency over NVIDIA’s Blackwell GPUs in specific benchmarks. The results are preliminary and vendor-reported, with deployment still pending.

OpenAI has published initial measured results for its Jalapeño inference chip, revealing significant efficiency and latency improvements over NVIDIA’s Blackwell systems in benchmark tests. These early results, though vendor-reported and not yet deployed, suggest that OpenAI’s custom silicon could reduce AI inference costs and improve responsiveness for large language models, making it a notable development in AI hardware.

OpenAI’s Jalapeño chip was tested against NVIDIA’s Blackwell generation on the InferenceX benchmark, which measures the full process of serving AI requests across three different models: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. The tests showed that Jalapeño achieved between 1.5 to 1.9 times higher performance per watt, and 1.7 to 3.6 times lower latency, compared to NVIDIA systems. These results were obtained on open models not owned by OpenAI, indicating the chip’s versatility beyond its proprietary models.

OpenAI emphasizes that these measurements are preliminary, vendor-reported, and based on a prototype chip not yet deployed in production. The company plans to begin integrating Jalapeño into its infrastructure by the end of 2024, pending further validation. The performance gains are notable, but the results are specific to inference tasks and are not directly comparable to general-purpose GPUs used for training or other workloads.

At a glance
reportWhen: announced March 2024
The developmentOpenAI’s Jalapeño chip has demonstrated promising inference performance metrics in early tests, highlighting its potential for more efficient AI serving.
AI DISPATCH · REALITY CHECKOpenAI Jalapeño · part 1 of 2 · 25 Aug 2026
The numbers are strong — and they’re the vendor’s
Jalapeño’s First Results: Read the Metric, Not the Headline

OpenAI’s first custom inference chip posts real per-watt wins on a public benchmark — measured by OpenAI, on the metric OpenAI chose, against NVIDIA only, on a chip not yet deployed.

1.5–1.9×
More AI work per watt (peak)
1.7–3.6×
Lower end-to-end latency
2.1–4.1×
Higher on interactive workloads
Per-watt inference — InferenceX (SemiAnalysis), OpenAI-run
Three external models, all vs NVIDIA Blackwell

Normalized by published TDP: Jalapeño 700W (measured ≤550W) vs GB200 1,200W / GB300 1,400W. Peak throughput per kW — higher is better.

GPT-OSS 120B mixed TPS / kW
vs GB200 · ~1.9×
Jalapeño
85.4k
GB200
45.0k
DeepSeek R1 670B mixed TPS / kW
vs GB300 · ~1.7×
Jalapeño
19.6k
GB300
11.8k
Kimi K2.5 1T mixed TPS / kW · largest tested
vs GB300 · ~1.5×
Jalapeño
18.2k
GB300
11.9k
Read the metric — three things the headline hides
~“Per watt” is a choice. Defensible for datacenter economics, but it structurally favors the lower-power part. Per-chip or per-dollar would read differently.
!ASIC vs general-purpose GPU. Blackwell trains and infers; Jalapeño does one job. Beating a GPU on inference-per-watt is why you build an ASIC — not a full verdict on the GPU.
iVendor-reported, not yet deployed. OpenAI’s own measurements; ships inside OpenAI by year-end, qualification ongoing. Ignore the 50–100× “at previous TBT” cherry — it’s one narrow operating point.

Performance Improvements Could Lower AI Costs

The reported efficiency and latency improvements suggest that Jalapeño could reduce the operational costs of large-scale AI inference, which is a major expense for data centers. By focusing on power efficiency and optimizing for inference workloads, OpenAI's hardware could enable faster, cheaper AI services, potentially impacting the economics of deploying large language models at scale.

However, these are early results, and the chip is not yet in production. The significance lies in the potential for hardware specialization to reshape AI infrastructure, but confirmation through independent benchmarks and real-world deployment remains pending.

Amazon

AI inference hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

OpenAI's Hardware Development and Benchmarking Background

OpenAI has been exploring custom hardware solutions to improve AI inference efficiency, recognizing that general-purpose GPUs are increasingly costly and power-hungry at scale. The company announced Jalapeño in early 2024, framing it as a dedicated inference ASIC designed around workload phases, with a focus on minimizing data movement and optimizing for the agentic nature of language models.

The initial performance figures come from OpenAI's own testing against NVIDIA's Blackwell chips, specifically targeting inference tasks using publicly available models on the InferenceX benchmark. Prior to Jalapeño, OpenAI relied on third-party hardware, but the new chip aims to reduce costs and improve performance for its own infrastructure and potentially for external clients.

While the numbers are promising, they are early and vendor-reported, and the chip has not yet been deployed operationally. Industry observers note that hardware tailored specifically for inference is an emerging trend, with other companies also investing in ASICs, but OpenAI's results are among the first to showcase dedicated inference hardware with such performance metrics.

Amazon

AI accelerator cards

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Aspects and Deployment Timeline

It remains unclear how Jalapeño will perform in real-world, large-scale deployments beyond the initial tests. The measurements are vendor-reported and have not been independently verified. Additionally, the chip is still in the qualification stage and is not yet deployed within OpenAI's infrastructure, leaving questions about its actual operational performance and reliability.

Further, it is not confirmed whether Jalapeño will be adopted widely outside OpenAI or how it compares to future generations of GPUs or other ASICs developed by competitors.

Amazon

custom AI inference chips

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Upcoming Deployment and Validation Milestones

OpenAI plans to complete the qualification process for Jalapeño and begin deploying the chip internally by late 2024. Independent benchmarking and third-party testing are expected to follow, which will clarify its performance and cost advantages. The company may also explore licensing or selling the hardware to external clients, but no official plans have been announced.

Further developments will include real-world performance data, long-term reliability assessments, and comparisons against other inference hardware options, shaping the future landscape of AI infrastructure.

Amazon

high performance AI GPUs

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Jalapeño, and why is it important?

Jalapeño is OpenAI's custom-designed inference chip aimed at improving AI request processing efficiency, potentially lowering costs and latency for large language models.

Are the performance results confirmed?

The results are vendor-reported, preliminary, and based on early prototypes. Independent validation and real-world deployment are still pending.

How does Jalapeño compare to NVIDIA GPUs?

In initial tests, Jalapeño shows higher efficiency and lower latency for inference tasks, but these are limited to specific benchmarks and not comprehensive comparisons against all GPU workloads.

When will Jalapeño be deployed in OpenAI's infrastructure?

OpenAI plans to begin internal deployment by the end of 2024, with further validation and testing to follow.

Could Jalapeño influence AI hardware development industry-wide?

If validated, Jalapeño's performance could encourage more specialization in inference hardware, potentially impacting the economics of large-scale AI deployment.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

Delvasta: Forms That Build Themselves

Delvasta introduces an early-access platform that automatically creates adaptive, branching forms, quizzes, and funnels to improve lead capture and data quality.

The European Union: Rules First, Cushion Always

The EU is prioritizing regulation and social protections over ownership in managing AI and labor transitions, with significant implications for workers and policy.

The Co-Founder’s Black Hole — A Structural Read on Jack Clark’s Automated AI R&D Essay

Jack Clark predicts over 60% chance of fully automated AI research by 2028, raising concerns about institutional readiness and future risks.

A Frontier AI Model Just Went Dark For 18 Days. The Kill-Switch Is Real Now.

A leading AI model was globally disabled for 18 days following US government orders, marking a new era of AI governance and control.