Mistral just doesn't seem like an interesting company to me any more.
They started off as the kind of people who released their models as magnet links and made them actually user-aligned instead of California-aligned. This is what I like to see from an AI company. Now, their models are no different from Open AI, Anthropic, Google, Meta and everybody else.
Could you expand more on what you mean by "California-aligned" (especially if you are thinking beyond AI models, though that alone could be interesting).
The way I see it, there are three kinds of models:
1. Unaligned models: You ask them a question, they complete your prompt with more details about the question to be answered instead of the answer. This is what you get from a basic LLM if you train it on the internet. It was how GPT-2 and GPT-3 worked before Chat GPT was invented. Such models aren't very useful for chat, you need to do weird prompt engineering tricks to actually get an answer instead of a clarification of your question. The original Mistral 7b was also of this kind.
2. User-aligned models. Aligned to answer questions and follow instructions, but no more and no less. If you ask them how to kill your wife, how to cook meth or how to make an atomic bomb, they'll happily help you and regurgitate the facts you can already find on the internet. They have no access to non-public information, the "dangerous" things they can tell you are already pretty easy to find, so the danger if you're using them for chat is actually minimal. However, they're far easier to use for troll farms, mass phishing campaigns, sockpuppets, political misinformation campaigns etc. If you ask them to engage with a pro-Biden tweet in the most triggering and overtly racist way possible and throw some pro-Putin angle into the mix, they'll happily accommodate your request, and you can do this at scale, paying orders of magnitude less than you would for a content farm. Mistral 7b instruct is a good example of such a model.
3. California-aligned models. They'll happily fulfill your request, unless it conflicts with DEI ideology, which has a particular foothold in CA, sort of exists in other parts of the US and is completely absent in non-English-speaking countries, even among extremely left-leaning populations. Google Gemini is the most egregious example.
It will be interesting to see if/how they walk back from releasing the weights. They have put a lot of effort into the "open" (not open source but very intentionally trying to conflate) approach to AI. But probably almost nobody is paying attention.
On the other hand, Meta has very little to gain by closing off their models. What would they do with them and who would use them? Llama was a coup because even though the license sucked, a passable model with available weights and llama.cpp allowed it to soar over the others of the time. Hopefully the benefit they got from that trumps any calls from the safety crowd to not share weights.
Really depends how good llama 3 will be. I see people thinking it would be better than GPT 4. But they still need to build a model as good as GPT 3.5 even, as llama 2 is... not.
Comments
Mistral just doesn't seem like an interesting company to me any more.
They started off as the kind of people who released their models as magnet links and made them actually user-aligned instead of California-aligned. This is what I like to see from an AI company. Now, their models are no different from Open AI, Anthropic, Google, Meta and everybody else.
Could you expand more on what you mean by "California-aligned" (especially if you are thinking beyond AI models, though that alone could be interesting).
I assume they mean "ethnically diverse Nazis"[1].
[1]: https://www.nytimes.com/2024/02/22/technology/google-gemini-...
As an aside its absolutely mad that the sole NYT takeaway from the whole debacle is putting people of color in Nazi uniforms.
This kind of framing is the NYT’s equivalent of clickbait for their subscribers. They also practice some clickbait as well, but imho not a lot.
(I’m a subscriber but this kind of thing irritates me, and makes the paper less useful, the equivalent of Google’s searches being less useful)
The way I see it, there are three kinds of models:
1. Unaligned models: You ask them a question, they complete your prompt with more details about the question to be answered instead of the answer. This is what you get from a basic LLM if you train it on the internet. It was how GPT-2 and GPT-3 worked before Chat GPT was invented. Such models aren't very useful for chat, you need to do weird prompt engineering tricks to actually get an answer instead of a clarification of your question. The original Mistral 7b was also of this kind.
2. User-aligned models. Aligned to answer questions and follow instructions, but no more and no less. If you ask them how to kill your wife, how to cook meth or how to make an atomic bomb, they'll happily help you and regurgitate the facts you can already find on the internet. They have no access to non-public information, the "dangerous" things they can tell you are already pretty easy to find, so the danger if you're using them for chat is actually minimal. However, they're far easier to use for troll farms, mass phishing campaigns, sockpuppets, political misinformation campaigns etc. If you ask them to engage with a pro-Biden tweet in the most triggering and overtly racist way possible and throw some pro-Putin angle into the mix, they'll happily accommodate your request, and you can do this at scale, paying orders of magnitude less than you would for a content farm. Mistral 7b instruct is a good example of such a model.
3. California-aligned models. They'll happily fulfill your request, unless it conflicts with DEI ideology, which has a particular foothold in CA, sort of exists in other parts of the US and is completely absent in non-English-speaking countries, even among extremely left-leaning populations. Google Gemini is the most egregious example.
Meta actually releases the weights though (for now)
It will be interesting to see if/how they walk back from releasing the weights. They have put a lot of effort into the "open" (not open source but very intentionally trying to conflate) approach to AI. But probably almost nobody is paying attention.
On the other hand, Meta has very little to gain by closing off their models. What would they do with them and who would use them? Llama was a coup because even though the license sucked, a passable model with available weights and llama.cpp allowed it to soar over the others of the time. Hopefully the benefit they got from that trumps any calls from the safety crowd to not share weights.
Really depends how good llama 3 will be. I see people thinking it would be better than GPT 4. But they still need to build a model as good as GPT 3.5 even, as llama 2 is... not.
But surely they're still remain as they were, i.e. not California aligned?