Comment on Choosing an AI model: one prompt, 11 models, different resultsComments−plumb_samji30dInteresting exploratory comparison, but I be cautious about treating it as a model benchmark With only three runs per model, the results are highly sensitive to randomness
Comments
Interesting exploratory comparison, but I be cautious about treating it as a model benchmark With only three runs per model, the results are highly sensitive to randomness