Comment on Qwen3-4B-Thinking-2507Comments−jampa1yI am reading this right, is this model way better than Gemma 3n[1]? (For only the benchmarks that are common among the models)=====LiveCodeBenchE4B IT: 13.2Qwen: 55.2===== AIME25E4B IT: 11.6Qwen: 81.3[1]: https://huggingface.co/google/gemma-3n-E4B−meatmanek1yReasoning models do a lot better at AIME than non-reasoning models, with o3 mini getting 85% and 4o-mini getting 11%. It makes some sense that this would apply to small models as well.
Comments
I am reading this right, is this model way better than Gemma 3n[1]? (For only the benchmarks that are common among the models)
=====
LiveCodeBench
E4B IT: 13.2
Qwen: 55.2
===== AIME25
E4B IT: 11.6
Qwen: 81.3
[1]: https://huggingface.co/google/gemma-3n-E4B
Reasoning models do a lot better at AIME than non-reasoning models, with o3 mini getting 85% and 4o-mini getting 11%. It makes some sense that this would apply to small models as well.