None of these models are the real Deepseek R1 that you can access via the API or chat!
The big one is a quantized version (it uses 4 bit per weight) and even that you probably cant run.
The other ones are fine-tunes of LLama 3.3 and Qwen2 which have been additionally trained on outputs of the big "Deepseek V3 + R1" model.
I'm happy people are looking into selfhosting models, but if you want to get an idea of what R1 can do, this is not a good way to do so.
Comments
None of these models are the real Deepseek R1 that you can access via the API or chat! The big one is a quantized version (it uses 4 bit per weight) and even that you probably cant run.
The other ones are fine-tunes of LLama 3.3 and Qwen2 which have been additionally trained on outputs of the big "Deepseek V3 + R1" model.
I'm happy people are looking into selfhosting models, but if you want to get an idea of what R1 can do, this is not a good way to do so.