Mistral seems to have given up on reaching the frontier and focuses instead on niches like moderation, OCR, and robot navigation.
That's not actually true. They at least claim they have a new attempt at a large model they plan to release near the end of this summer.
Their last 'large' model was a flop that needed to be replaced by their Medium 3.5, but they have not yet given up on large models, even if they're also making smaller more specialized models for now (because that's somewhere they actually do rather well).
I considered mentioning that in the footnote but my own skepticism stopped me. I'm now thinking I need to see it to believe it in terms of Mistral training a frontier model.
That's not to say I don't hope they will. It would be incredible if they could do that despite their limited funding and the wide swath of projects they are working on. I do think that lack of focus is going to be a big problem.
Fair enough, though I think it's rather unlikely that they'd be lying about their work on a big model, even if one might be reasonably skeptical if the model will be any good.
I think it's rather important for them that they at least keep attempting to build frontier models, because if they give up on that, then governments will no longer see them as a strateigic asset and instead just see them as a regular business.
I think that for better or worse, Mistral wants the buy-in from governments to scale up, and even if they have faltering outcomes, it's better to have that, and be able to say "look, this is what we tried, we failed due to lack of compute. We need more money for more compute."
Oh definitely, not saying they are lying about working on a big model, but as you say I'm skeptical it will be any good.
I hope they will get that boost, but I'm afraid that they haven't built credibility so far to convince governments to buy in at the level needed. Their slip away from the frontier might even have made governments more cautious on investing in AI development, which is the opposite of what we need right now.
Comments
Regarding
That's not actually true. They at least claim they have a new attempt at a large model they plan to release near the end of this summer.
Their last 'large' model was a flop that needed to be replaced by their Medium 3.5, but they have not yet given up on large models, even if they're also making smaller more specialized models for now (because that's somewhere they actually do rather well).
I considered mentioning that in the footnote but my own skepticism stopped me. I'm now thinking I need to see it to believe it in terms of Mistral training a frontier model.
That's not to say I don't hope they will. It would be incredible if they could do that despite their limited funding and the wide swath of projects they are working on. I do think that lack of focus is going to be a big problem.
Fair enough, though I think it's rather unlikely that they'd be lying about their work on a big model, even if one might be reasonably skeptical if the model will be any good.
I think it's rather important for them that they at least keep attempting to build frontier models, because if they give up on that, then governments will no longer see them as a strateigic asset and instead just see them as a regular business.
I think that for better or worse, Mistral wants the buy-in from governments to scale up, and even if they have faltering outcomes, it's better to have that, and be able to say "look, this is what we tried, we failed due to lack of compute. We need more money for more compute."
Oh definitely, not saying they are lying about working on a big model, but as you say I'm skeptical it will be any good.
I hope they will get that boost, but I'm afraid that they haven't built credibility so far to convince governments to buy in at the level needed. Their slip away from the frontier might even have made governments more cautious on investing in AI development, which is the opposite of what we need right now.