Loading knowledge network

Patman's Neural Network

AI & Technology

OpenAI bets 10 billion on speed

Why the partnership with Cerebras turns the AI industry upside down

Published on 15 January 2026

Translated from German

OpenAI has just announced a deal that makes the AI world take notice. Over 10 billion dollars for Cerebras chips. A three-year term. A clear signal: inference is the new gold, and speed is king.

The story actually begins with Google. In November 2024, Google launched Gemini 3. A frontier model that dominated the benchmarks. But the special thing: it ran entirely on TPUs, not on Nvidia GPUs. Training and inference, all on Google's own chips.

Nvidia quickly realized: specialized chips are a real threat. The response followed at Christmas 2025. Nvidia "bought" Groq for 20 billion dollars. With "bought" in quotation marks here. Practically the entire team moved to Nvidia, they licensed the technology, but the company Groq continues to exist. A maneuver to circumvent antitrust authorities.

Then came January 14, 2026. OpenAI announces the partnership with Cerebras. 750 megawatts of computing power over three years. According to the Wall Street Journal: worth over 10 billion dollars.


Why Cerebras?

Sam Altman faced a problem: too much dependency on Nvidia. All GPUs from there, all inference partners use Nvidia chips. And Groq? Just sold to Nvidia. The platform risk was too big.

Cerebras offered the perfect alternative. The chips are different. Radically different. Instead of connecting hundreds of small chips, they use an entire wafer as a single processor. The result: up to 15 times faster responses than GPU systems.

Speed

Cerebras delivers over 3,000 tokens per second. Groq manages 465. GPT-OSS on Cerebras: 2,700+ tokens/second vs. 900 on Nvidia's Blackwell B200. That's not incrementally better. That's a different league.

Architecture

The secret lies in SRAM. Cerebras packs 44 GB directly onto the chip. No external memory like with GPUs. That means: no latency during data transfer. Everything is right where it's needed.

Independence

While GPU manufacturers suffer from memory shortages and prices skyrocket, Cerebras is unaffected. CEO Andrew Feldman confirmed: "We don't use that. That's an advantage for us."

Costs

Cerebras is 32% cheaper than Nvidia's Blackwell B200. At 21 times the speed. That results in massive price-performance leadership.


Inference is the new training

The industry has understood: the money is not in training. Training is one-time. It costs a lot, but then the model is ready. Inference runs forever. The more users, the more revenue. Training is a cost center. Inference is the revenue generator.

And here, speed becomes the deciding factor. For coding agents, it makes the difference between flow and frustration. For reasoning models like o1 or DeepSeek R1, which generate thousands of "thinking tokens," the wait time shrinks from minutes to seconds.

Through Cerebras, OpenAI suddenly gets massive additional capacity. That means: all Nvidia GPUs can now be used for training. No longer for inference. The result: better models in less time.


What this means for the future

Specialized chips are winning. Nvidia remains dominant in training, but inference belongs to specialized players. Cerebras, TPUs, and others will divide the market among themselves.

Cerebras is presumably on the verge of an IPO. They already filed documents in 2024, then withdrew them, then raised funding again. This OpenAI deal is the final boost. Every frontier lab is watching closely now.

And OpenAI? They just solved their biggest problem. Their only limiting factor was capacity. Now they have more of it. From a different source. With better performance. At lower costs.

The race for compute escalates

We are witnessing an all-out war for computing capacity. Google uses TPUs. Nvidia buys Groq. OpenAI partners with Cerebras. The message is clear: whoever wants to win in the AI era needs the fastest chips. Not the biggest. The fastest.

Return to network