Comment on LLM Benchmark: Frontier models now statistically indistinguishableComments−anonzzzies8moWould be nice to include similar sized open (source/weights) ones.−js4everOP8moJust tried devstral 2 (123B from Mistral) it scored 76% ... Disappointing
Comments
Would be nice to include similar sized open (source/weights) ones.
Just tried devstral 2 (123B from Mistral) it scored 76% ... Disappointing