Comment on DeepSeek-V3Comments−janice19991ya strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token.What kind of hardware do you need to run this?−bavell1y8x H200s recommended:https://github.com/sgl-project/sglang/tree/main/benchmark/de...−boroboro41yThey discuss it in the paper and recommend 32 GPUs (H800 in their case) for prefill stage and 320 GPUs for decoding.=)
Comments
What kind of hardware do you need to run this?
8x H200s recommended:
https://github.com/sgl-project/sglang/tree/main/benchmark/de...
They discuss it in the paper and recommend 32 GPUs (H800 in their case) for prefill stage and 320 GPUs for decoding.
=)