All posts

Infrastructure2 min read

Colossus 2 aims to more than double its GPUs by year-end

Elon Musk laid out the most detailed timetable yet for xAI's Memphis supercomputer: from 550,000 Nvidia chips today to about 1.21 million by late December.

Rows of gradient bars rising from a baseline on a deep indigo background

The compute race got a new set of numbers this week. In a post on X on September 25, Elon Musk said Colossus 2, xAI's AI computing cluster in the Memphis area, could more than double its Nvidia chip count by the end of 2026. Bloomberg called it the most detailed timetable yet for the expansion.

The numbers

By Musk's account, Colossus 2 currently runs:

  • 110,000 Nvidia GB200 chips
  • 440,000 Nvidia GB300 chips

He then laid out three more batches of 220,000 GB300s each: one going live next week, another in November, and a third in late December, "if we get lucky."

Stacked bar chart of Colossus 2 chip counts: 550K today, 770K next week, 990K in November and 1.21M in late December
Chip counts by Musk's stated timetable. The December step depends on things going well.

Adding 660,000 chips would take the facility from about 550,000 to about 1.21 million, more than double. The company had previously said it planned to equip the Memphis facility with 1 million GPUs in 2026.

Why it matters

xAI has made the Colossus build-out central to its effort to close the compute gap with OpenAI, Google and Anthropic. Buying chips is only part of the problem. Racks, power, cooling and networking at this scale are what decide how fast new hardware actually starts training models.

Musk's timelines don't always hold, and the December batch is explicitly conditional. Still, the direction matches the rest of the industry: frontier labs are measuring training clusters in the millions of accelerators.

What it means for teams building on AI

Most companies will never touch a GB300. But frontier infrastructure spending still reaches the teams building on top of it:

  1. More compute usually means cheaper tokens. This month's price war between Anthropic and OpenAI is partly what happens as capacity like this comes online. Plan for per-token prices to keep falling.
  2. Launches get more frequent. Bigger clusters shorten training runs, which means more frequent model updates. A vendor-neutral evaluation harness will pay off many times over.
  3. Capacity limits are easing. Rate limits and waitlists have held back many agent deployments. As supply catches up, the constraint moves from "can we get tokens" to "can we use them well."

The bottom line

Numbers like 1.21 million GPUs are hard to picture, but the practical effect is simple: frontier intelligence is becoming more abundant. The advantage shifts to teams that can turn cheap intelligence into dependable workflows.

Keep reading

All posts ↗