Model Advisor

Big open-source models, and the least hardware that will run them. We show the memory arithmetic out loud so you can see where the number comes from.

License checked against each model's official card. We only list models under MIT or Apache 2.0 - the two most permissive open-source licenses.

The rule we use: a model needs roughly 1 GB of GPU memory per billion parameters, plus 20% working room. Then we buy the fewest server GPUs whose memory adds up to at least that.

DeepSeek-R1

MIT 671B total params 37B active

A frontier open reasoning model. Mixture-of-Experts: 671B total parameters, ~37B active per token.

Memory math
671B Γ— 1 GB = 671 GB, +20% = 806 GB minimum
Minimum memory to run it
806 GB

The least fast memory a setup needs before this model will even load and run. Less than this and the model simply will not fit.

Minimum setup you must buy

Combined memory
864 GB

All the GPU memory in this build added together.

Combined power
4200 W

Every machine in this build drawing electricity at the same time.

Total price
$150,000

The price of everything in this build added up.

⚑

Draws as much electricity as 3.5 average homes running around the clock.

Left running, it would drain a full 90 kWh electric-car battery every 21.4 hours.

πŸ“¦ This build is a cluster

A cluster is multiple machines wired together to act as one bigger computer.

Be aware: real datacenter clusters use very high-speed links (NVLink / NVSwitch inside a box, InfiniBand between boxes) so the GPUs pool their memory and behave like one huge GPU. A pile of desktop cards on ordinary PCIe slots and consumer networking is not the same - the connections are far slower, the memory does not pool cleanly, and you cannot simply stack desktop cards to run one enormous model efficiently. For serious large-model work, server GPUs and real interconnects are worth the money. Read the honest comparison β†’

Ready to buy this build?

Add everything above to your cart in one click, then check out with just your name and a way to reach you.

GLM-4.6

MIT 355B total params 32B active

A strong open agentic model from Z.ai. Mixture-of-Experts: 355B total, ~32B active per token.

Memory math
355B Γ— 1 GB = 355 GB, +20% = 426 GB minimum
Minimum memory to run it
426 GB

The least fast memory a setup needs before this model will even load and run. Less than this and the model simply will not fit.

Minimum setup you must buy

Combined memory
576 GB

All the GPU memory in this build added together.

Combined power
2800 W

Every machine in this build drawing electricity at the same time.

Total price
$100,000

The price of everything in this build added up.

⚑

Draws as much electricity as 2.3 average homes running around the clock.

Left running, it would drain a full 90 kWh electric-car battery every 32.1 hours.

πŸ“¦ This build is a cluster

A cluster is multiple machines wired together to act as one bigger computer.

Be aware: real datacenter clusters use very high-speed links (NVLink / NVSwitch inside a box, InfiniBand between boxes) so the GPUs pool their memory and behave like one huge GPU. A pile of desktop cards on ordinary PCIe slots and consumer networking is not the same - the connections are far slower, the memory does not pool cleanly, and you cannot simply stack desktop cards to run one enormous model efficiently. For serious large-model work, server GPUs and real interconnects are worth the money. Read the honest comparison β†’

Ready to buy this build?

Add everything above to your cart in one click, then check out with just your name and a way to reach you.

Qwen3-235B-A22B

Apache 2.0 235B total params 22B active

Alibaba's top open model. Mixture-of-Experts: 235B total parameters, ~22B active per token.

Memory math
235B Γ— 1 GB = 235 GB, +20% = 282 GB minimum
Minimum memory to run it
282 GB

The least fast memory a setup needs before this model will even load and run. Less than this and the model simply will not fit.

Minimum setup you must buy

Combined memory
288 GB

All the GPU memory in this build added together.

Combined power
1400 W

Every machine in this build drawing electricity at the same time.

Total price
$50,000

The price of everything in this build added up.

⚑

Draws as much electricity as 1.2 average homes running around the clock.

Left running, it would drain a full 90 kWh electric-car battery every 64.3 hours.

Ready to buy this build?

Add everything above to your cart in one click, then check out with just your name and a way to reach you.

Mixtral 8x22B

Apache 2.0 141B total params 39B active

Mistral AI's largest open-weight model. Mixture-of-Experts: 141B total, ~39B active per token.

Memory math
141B Γ— 1 GB = 141 GB, +20% = 170 GB minimum
Minimum memory to run it
170 GB

The least fast memory a setup needs before this model will even load and run. Less than this and the model simply will not fit.

Minimum setup you must buy

Combined memory
180 GB

All the GPU memory in this build added together.

Combined power
1000 W

Every machine in this build drawing electricity at the same time.

Total price
$40,000

The price of everything in this build added up.

⚑

Draws as much electricity as 0.8 average homes running around the clock.

Left running, it would drain a full 90 kWh electric-car battery every 90.0 hours.

Ready to buy this build?

Add everything above to your cart in one click, then check out with just your name and a way to reach you.