According to this, you can fit it on a CPU-only setup (no GPUs) with 2 x AMD EPYC CPUs and 24 x 32GB DDR5-RDIMM RAM. About $6000 MSRP for the rig. Doubt you are going to get very many tokens/sec out of it though (6-8, according to the author).
The minimum deployment unit of the decoding stage consists of 40 nodes with 320 GPUs.
but realistically, >=671GB of VRAM to run at full precision on GPU, or >=131G VRAM to run the most heavily quantized version[1], or >=671GB RAM, and a dose of patience to run on CPU.
Comments
So what hardware do I need to run DeepSeek R1 with 670B parameters?
https://www.reddit.com/r/LocalLLaMA/comments/1ic8cjf/6000_co...
According to this, you can fit it on a CPU-only setup (no GPUs) with 2 x AMD EPYC CPUs and 24 x 32GB DDR5-RDIMM RAM. About $6000 MSRP for the rig. Doubt you are going to get very many tokens/sec out of it though (6-8, according to the author).
Pretty impressive TPS numbers for CPU-only
Per the technical report:
but realistically, >=671GB of VRAM to run at full precision on GPU, or >=131G VRAM to run the most heavily quantized version[1], or >=671GB RAM, and a dose of patience to run on CPU.
[1]: https://news.ycombinator.com/item?id=42850222