All posts

Hardware

OpenAI and Broadcom unveil Jalapeño, OpenAI's first custom inference chip

OpenAI designed it, Broadcom built it, and the companies say it went from design to tape-out in nine months. No specs or benchmarks yet; deployment starts late 2026.

HackHoster Team · · 11 min read

Close-up of a silicon wafer covered in a grid of identical chip dies, catching rainbow reflections from a bright light
Photo: Le hollandais volant / Wikimedia Commons, CC BY 4.0

At a glance

  • OpenAI and Broadcom unveiled Jalapeño on June 24, 2026, an accelerator built only for running large language models, not training them.
  • The companies say the chip went from first design to tape-out in about nine months; Tom's Hardware puts a typical new ASIC at 1.5 to 2 years.
  • Engineering samples run in OpenAI's lab at target clock and power, including GPT-5.3-Codex-Spark, which today runs on Cerebras hardware.
  • No compute, memory, power or benchmark figures were released; Tom's Hardware estimates the compute die at about 840 square millimeters from photos.
  • Broadcom CEO Hock Tan told CNBC deployment starts small in late 2026, ramps in 2027 and goes full tilt in the first half of 2028.

OpenAI and Broadcom on Wednesday showed Jalapeño, the first chip OpenAI has designed itself. It is an inference accelerator, meaning it runs models that are already trained rather than training new ones, and OpenAI says it was built from scratch for serving large language models rather than adapted from a general-purpose AI processor. The companies call it the first Intelligence Processor in a platform they plan to extend over several generations.

According to The Decoder, Broadcom CEO Hock Tan and president Charlie Kawwas handed the first wafer to Sam Altman and Greg Brockman at the reveal. The companies say engineering samples are already running in OpenAI's labs at the target clock speed and power, and OpenAI says they are running real machine-learning workloads, including GPT-5.3-Codex-Spark, a fast coding model that currently runs on Cerebras hardware. Initial deployment is planned by the end of 2026.

What the companies did not release is almost everything a hardware engineer would ask for first: compute throughput, memory capacity and bandwidth, power draw, process node and benchmark results. This article covers what is known, what can be inferred, and what is still marketing.

Sam Altman, in a grey sweater, seated on a stage with a blurred yellow and green backdrop
OpenAI CEO Sam Altman at TED in April 2025. He and Greg Brockman received the first Jalapeño wafer from Broadcom's leadership. Photo: Steve Jurvetson / Wikimedia Commons, CC BY 2.0

Why OpenAI wants its own chip

OpenAI has been one of the largest buyers of Nvidia GPUs since ChatGPT launched in 2022, and CNBC reports that demand for its products has grown so fast that it needs other sources of silicon. Brockman told CNBC the company cannot get compute fast enough. Over the past year OpenAI has signed compute agreements with Nvidia, AMD, Cerebras and Amazon Web Services, the last of which includes Amazon's Trainium chips.

The Register's coverage of the October 2025 Broadcom announcement lists the scale of those commitments at the time: a 10-gigawatt agreement with Nvidia that came with up to $100 billion of Nvidia investment, a 6-gigawatt agreement with AMD that included a warrant for up to 160 million AMD shares, and the 10-gigawatt Broadcom plan. The Register noted that, unlike the Nvidia and AMD deals, the Broadcom arrangement did not involve either company investing in the other.

Building a custom chip is the route Google took a decade earlier. According to Wikipedia's history of the Tensor Processing Unit, Google began using TPUs inside its data centers in 2015 and announced them at Google I/O in May 2016. The first TPU was an inference-only chip built around 8-bit arithmetic. Broadcom has co-developed Google's TPUs, turning Google's specifications into manufacturable silicon, and TechCrunch frames Jalapeño as OpenAI following Google and Amazon down the same road.

A Google TPU v4 circuit board with four liquid-cooled chip packages connected by colored coolant hoses
A Google TPU v4 board with four liquid-cooled packages, each an ASIC surrounded by high-bandwidth memory. Google's TPUs, co-developed with Broadcom, are the best-known precedent for Jalapeño. Image: Norman P. Jouppi et al. / Wikimedia Commons, CC BY 4.0

The people are part of that story too. In October 2024, Reuters and Bloomberg reported that OpenAI had assembled a chip team of about 20 people, led by Thomas Norrie and Richard Ho, both of whom had previously built TPUs at Google. The reports said OpenAI was working with Broadcom on an inference chip, planned to manufacture it at TSMC, and had dropped earlier, more ambitious plans to build its own network of chip factories because of the cost and time involved. Ho now leads OpenAI's hardware program and is quoted in this week's announcement.

Who built what

The division of labor follows a common pattern for custom AI chips, in which the company that runs the workload specifies the chip and an experienced partner turns the design into manufacturable silicon. According to The Decoder:

  • OpenAI defined the architecture, based on what it knows about how its models behave in production.
  • Broadcom did the silicon implementation and supplies networking, including its Tomahawk switch chips.
  • Celestica handles board, rack and system-level integration.
The Broadcom headquarters in San Jose, California: a low office building behind a row of large trees, with a red and white Broadcom sign by the road
Broadcom's headquarters in San Jose, California. Broadcom turned OpenAI's architecture into silicon and supplies the networking. Photo: Coolcaesar / Wikimedia Commons, CC BY-SA 4.0

CNBC reports that OpenAI and Broadcom had worked together for 18 months before announcing their partnership in October 2025, and that OpenAI also designed large parts of the computer system the chip will sit in. TechCrunch describes OpenAI's ambition as owning the whole stack: chip architecture, kernels, memory systems, networking, scheduling and deployment. Brockman's rationale, as quoted by TechCrunch: "We have a deep understanding of the workload."

There is also a commercial detail worth knowing. The Decoder, citing earlier reporting, says Broadcom asked Microsoft to guarantee it would buy 40 percent of the chips to secure the first phase. In the announcement, Hock Tan said the chip will enable gigawatt-scale data centers with Microsoft and other partners, as quoted by Tom's Hardware.

For Broadcom, OpenAI joins a short list of very large custom-chip customers. Tan told CNBC that compute demand from the company's six customers is far more than Broadcom can address, and that he sees it continuing into 2028. CNBC notes that Broadcom shares were up 10% for 2026 before the announcement and have grown almost sevenfold since the end of 2022. When the partnership was first announced in October 2025, The Register reported that Broadcom would supply Ethernet, PCIe and optical connectivity for complete systems aimed mainly at inference.

How an inference chip differs from a GPU

Training and inference stress hardware differently. Training pushes huge batches of data through a model and updates every weight, which rewards raw arithmetic. Serving a chat or coding model is different: the model generates one token at a time, and for each token the chip has to read a large share of the model's weights from memory. That makes inference heavily dependent on memory bandwidth and on how far data has to travel inside the system.

OpenAI's description of Jalapeño targets exactly those costs. According to Tom's Hardware, the companies say the design addresses expensive data movement, the balance between compute and memory, and networking efficiency, and that it is built to run close to its theoretical peak rather than leaving hardware idle. Richard Ho, who leads OpenAI's hardware program, said in the announcement, as quoted by Tom's Hardware, that the team optimized the architecture around the kernels, memory movement, networking and serving patterns that matter most for frontier models.

ASIC, defined. An application-specific integrated circuit is a chip designed for one job. CNBC notes that industry experts see ASICs as less flexible than Nvidia's GPUs, but cheaper and easier to tune for a specific AI task. The bet is that the job, serving transformer models, is stable enough to justify fixed silicon.

What the photos suggest

Because there are no specifications, Tom's Hardware's Anton Shilov worked from the wafer and package photos the companies released. His reading:

  • The package appears to hold one large compute chiplet surrounded by six HBM memory stacks, plus a second chiplet that likely handles input and output, flanked by two structural dummy dies.
  • Using the known size of an HBM package as a ruler, he estimates the compute die at about 25.5 by 33 millimeters, or roughly 840 square millimeters. That is close to the 858-square-millimeter maximum an EUV lithography machine can expose in one shot, the so-called reticle limit.
  • The die has a highly regular, repeated floorplan. It could be a systolic-array design like Google's TPUs, but he stresses that the image is not clean enough to say.

Tom's Hardware also notes that the choice of HBM, rather than the cheaper DRAM many inference accelerators use, points to a design that wants both high throughput and low latency. That fits OpenAI's push into reasoning models and agents, where a slow response is a worse product.

The Cerebras system that serves Codex-Spark today takes a different route to the same problem. Cerebras says its wafer-scale engine has the largest on-chip memory of any AI processor, which keeps data next to the compute instead of in external memory. Jalapeño, judging by its package, relies on a reticle-sized die fed by stacked HBM, an approach closer to GPUs and TPUs.

A diagram of a 2-by-2 weights-stationary systolic array over several clock cycles, with input values flowing down and partial sums flowing right through a grid of cells
How a weights-stationary systolic array multiplies matrices: weights sit in a grid of cells while inputs and partial sums pulse through. Google's TPUs use this design; whether Jalapeño does is not known. Diagram: EMJzero / Wikimedia Commons, CC0

The appeal of a systolic array for inference is that it reuses data. Each weight is loaded into a cell once, and many inputs flow past it, so the chip does more arithmetic per byte fetched from memory. Wikipedia's TPU article notes that the first-generation TPU used a 256-by-256 array of this kind.

The nine-month claim

The headline claim is speed of development. The companies say Jalapeño went from initial design to tape-out in about nine months, and Tom's Hardware puts the usual time to design an ASIC from scratch at 1.5 to 2 years.

Tape-out is the point where a finished chip design is handed off to the factory to make photomasks and start manufacturing. It is a milestone, not a product. Bring-up, yield improvement, system integration and software all come after it.

Two things explain the gap, and only one of them is new. Brockman told CNBC that the chip was designed end to end in nine months with help from OpenAI's own models, and said the degree of acceleration surprised the company. Neither company said which parts of the design flow the models handled. Tom's Hardware points to the other factor: Broadcom reuses a great deal of its logic across the custom chips it builds for different customers, which shortens any project it takes on.

The TSMC Fab 18 complex in the Southern Taiwan Science Park, large white factory buildings under a blue sky
TSMC Fab 18 in the Southern Taiwan Science Park. 2024 reports said OpenAI's first chip would be manufactured at TSMC; the companies have not said which fab or process node Jalapeño uses. Photo: 4300streetcar / Wikimedia Commons, CC BY 4.0

The numbers, such as they are

What is public comes from the companies, from Tom's Hardware's photo analysis, and from Hock Tan's comments to CNBC.

ItemWhat is knownSource
PurposeInference for large language models, designed to run any LLMOpenAI, via Tom's Hardware
Design to tape-outAbout 9 monthsOpenAI and Broadcom
Compute dieAbout 840 mm², near the 858 mm² reticle limitTom's Hardware estimate from photos
MemorySix HBM stacks visible on the packageTom's Hardware estimate from photos
EfficiencyPerformance per watt substantially better than current hardwareOpenAI and Broadcom claim, no figures
Compute, bandwidth, power, nodeNot disclosed
NetworkingBroadcom, including Tomahawk switchesThe Decoder
SystemsBoards and racks by CelesticaThe Decoder

The timeline is clearer than the specs:

DateMilestone
2015Google starts using TPUs in its data centers
October 2024Reports describe OpenAI's 20-person chip team working with Broadcom and TSMC
October 13, 2025OpenAI and Broadcom announce a 10-gigawatt custom accelerator plan
February 12, 2026GPT-5.3-Codex-Spark launches on Cerebras hardware
June 24, 2026Jalapeño unveiled; engineering samples running in OpenAI's lab
Late 2026Small prototype deployment, according to Hock Tan
2027Ramp-up, according to Tan
First half of 2028Full-scale deployment, according to Tan
A large silver Cisco packet-processing chip mounted on a green circuit board inside a network switch
A packet-processing ASIC inside an Ethernet switch, made by Cisco. Broadcom's Tomahawk chips do this switching job in many data centers and will link Jalapeño systems together. Photo: Pokiiri / Wikimedia Commons, CC BY-SA 4.0

What is still missing

The efficiency comparison has a moving target. As Tom's Hardware points out, beating today's Nvidia Blackwell and AMD MI350 parts is not the same as beating Nvidia's Rubin and AMD's MI400 parts, the generation Jalapeño will actually compete with once it ships in volume. The companies also did not say which chips they mean by current state of the art.

Key caveat. Every performance claim here is self-reported and unverified. The Decoder notes the numbers have not been finalized and that a technical report is supposed to follow. Until then, there is nothing to compare.

The schedule carries risk too. Tan's own timeline has volume arriving in 2027 and 2028, not this year. After tape-out come silicon bring-up, yield ramps, system integration and the software stack that schedules models onto the chip, and any of those can slip. Broadcom's own CEO said back in 2024, according to the Reuters report, that custom chips are not an easy product for any customer to deploy.

Then there is the question of who gets to use it. OpenAI says the chip is designed to run any LLM, not only its own. Tom's Hardware reads that as a possible opening to sell the hardware to others, if OpenAI can secure enough supply from Broadcom and TSMC, but notes it is unclear whether outside tenants will ever get access.

Finally, the efficiency case rests on a bet that today's model architectures stay stable. A GPU can run whatever researchers invent next, while an ASIC tuned for today's transformer serving patterns may suit a different kind of model less well. Hock Tan called Jalapeño the beginning of a multi-generation roadmap, and later generations are where any shift in model design would have to be absorbed.

What it means for builders

For developers building on OpenAI's API, nothing changes today. If Jalapeño delivers what OpenAI claims, the effect would show up later as cheaper or faster inference without any work on your side. TechCrunch reports that low operating cost for real-time coding models was a design priority, which makes coding tools the first place to look for changes.

A few practical points for teams planning products around AI inference:

  • Expect speed tiers. Codex-Spark already shows OpenAI serving the same product family on different hardware for different latency needs. Cerebras says the model runs at more than 1,000 tokens per second on its wafer-scale engine. Design your app so a faster, cheaper tier can be swapped in for interactive steps while slower models handle long background work.
  • Measure cost per task, not per token. Hardware changes tend to show up as new prices and new rate limits. Log tokens, latency and retries per completed task so you can tell quickly whether a price change helps you.
  • Keep the provider layer thin. Google and Amazon already run AI workloads on their own chips, and OpenAI is now following. An abstraction that lets you switch models and endpoints protects you from capacity crunches on any one of them.

For hardware-minded teams, the public material on Google's TPUs is the best open window into this kind of design, and systolic arrays are simple enough to simulate in an afternoon. Building a small matrix-multiply array in an FPGA or a cycle-level simulator is a good way to feel why data movement, not arithmetic, dominates inference cost.

A large Google data center in The Dalles, Oregon: long low industrial buildings with cooling towers beside a road
Google's data center in The Dalles, Oregon, not an OpenAI site. OpenAI and Broadcom's plan calls for 10 gigawatts of capacity built around OpenAI-designed accelerators. Photo: Tony Webster / Wikimedia Commons, CC BY 2.0

What to watch

As of June 25, the next concrete items are the technical report The Decoder says will follow, and with it the first hard numbers for compute, memory and power. After that, Tan's timeline sets the checkpoints: small prototype deployments late this year, a ramp in 2027, and full scale in the first half of 2028. The other signal to watch is whether any of OpenAI's existing workloads, Codex-Spark included, move from partner hardware onto Jalapeño, and whether OpenAI opens the chip to anyone outside its own data centers.

Sources