Comment on LLM Benchmark: Frontier models now statistically indistinguishableparentComments−js4everOP8moJust tried devstral 2 (123B from Mistral) it scored 76% ... Disappointing
Comments
Just tried devstral 2 (123B from Mistral) it scored 76% ... Disappointing