Skip to content
Industrymedium signalverified

OpenAI's first chip is measured per watt, and the watt is the part to read twice

Jalapeño delivers 1.5 to 1.9 times more work per watt and up to 3.6 times lower latency than unnamed comparison systems, OpenAI says. The chip is rated at 700 watts and drew at most 550 in the tests, while rivals were normalized by their published rating.

By Redakcija WebAiRadarPublished 2 min readwritten by a model
Image: OpenAI

OpenAI published the first measured results for Jalapeño, its own inference chip, on August 25. The company tested it on InferenceX, a public benchmark from SemiAnalysis that covers the whole path of serving a request, and ran three open models on it: GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T.

What OpenAI reports

Across the three models, OpenAI says Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. For highly interactive workloads — the case that matters for agents, where delays compound step after step — it reports 2.1 to 4.1 times higher performance. On Kimi K2.5, the largest model tested, the figures are roughly 1.5 times the performance per watt and 3.4 times lower latency.

The company frames throughput and latency as a tradeoff that existing hardware has to make and this architecture does not, and says internal testing on OpenAI's own frontier models widens the advantage further. Those internal numbers are not published.

The metric is doing work

OpenAI states its method plainly: results are normalized using each accelerator's published chip power rating, and it argues performance per unit of power is more useful than performance per chip. In the same paragraph it notes Jalapeño is rated at 700 watts but stayed at or below 550 watts on the tested workloads.

Those two sentences pull in opposite directions. If every system is divided by its nameplate figure, a chip that runs a fifth below its own rating is credited with efficiency it did not have to demonstrate — and a rival that runs at its rating gets no such discount. The comparison systems are also not named, so no one outside can rerun the same pairing.

  • Benchmark: InferenceX, published by SemiAnalysis — a public benchmark, not a private harness.
  • Models: GPT-OSS 120B, DeepSeek R1 670B, Kimi K2.5 1T — all open weights.
  • Normalization: each accelerator's published power rating, not measured draw.
  • Comparison systems: described as leading commercially available systems, not named.

Why it still matters

A company that designs the model, the serving software and now the silicon can tune all three against its own traffic, and that is a structural advantage no benchmark table captures. If the ratios hold at scale, the effect shows up where customers can feel it: response times under load, and the price of tokens over time.

For now this is a vendor publishing its own first results on hardware nobody else can buy. The useful move is to note the claim, note the method, and wait for InferenceX runs somebody else performed.

Sources

BrandsChatGPT

Related

Industrystrong signal

The 19 percent everyone is quoting does not mean what the headlines say

Stanford's August 2026 update finds young workers in AI-exposed jobs 19 percent below where they would be had hiring kept pace. The comparison to last year's 13 percent is between two different measures — like for like, it is 15 to 19.

Stanford Digital Economy Labverified

EU AI ACT50the article in force since August 2, 2026
Industrystrong signal

How to label AI content under Article 50, and which part of it is not your job

Article 50 of the EU AI Act has applied since August 2, 2026, and it binds anyone serving people in the Union, wherever the server is. Most of the panic is about the machine-readable marking requirement, which for a site owner who calls somebody else's API is somebody else's obligation. Here is what is actually yours: a chatbot that says what it is, published text that either carries a name or carries a label, and a deepfake that admits it.

Evropska komisijaverified

PEW RESEARCH35%of pages written since ChatGPT
Industrystrong signal

Pew puts a number on it: 35% of pages published after ChatGPT show AI authorship

Pew Research ran nearly half a million pages from Common Crawl through a detector. Among pages published after November 2022, more than a third came back with significant signs of AI writing, and .com domains showed it at ten times the rate of .edu and .gov.

TechCrunchverified