The way I see it, there are three kinds of models:
1. Unaligned models: You ask them a question, they complete your prompt with more details about the question to be answered instead of the answer. This is what you get from a basic LLM if you train it on the internet. It was how GPT-2 and GPT-3 worked before Chat GPT was invented. Such models aren't very useful for chat, you need to do weird prompt engineering tricks to actually get an answer instead of a clarification of your question. The original Mistral 7b was also of this kind.
2. User-aligned models. Aligned to answer questions and follow instructions, but no more and no less. If you ask them how to kill your wife, how to cook meth or how to make an atomic bomb, they'll happily help you and regurgitate the facts you can already find on the internet. They have no access to non-public information, the "dangerous" things they can tell you are already pretty easy to find, so the danger if you're using them for chat is actually minimal. However, they're far easier to use for troll farms, mass phishing campaigns, sockpuppets, political misinformation campaigns etc. If you ask them to engage with a pro-Biden tweet in the most triggering and overtly racist way possible and throw some pro-Putin angle into the mix, they'll happily accommodate your request, and you can do this at scale, paying orders of magnitude less than you would for a content farm. Mistral 7b instruct is a good example of such a model.
3. California-aligned models. They'll happily fulfill your request, unless it conflicts with DEI ideology, which has a particular foothold in CA, sort of exists in other parts of the US and is completely absent in non-English-speaking countries, even among extremely left-leaning populations. Google Gemini is the most egregious example.
Comments
The way I see it, there are three kinds of models:
1. Unaligned models: You ask them a question, they complete your prompt with more details about the question to be answered instead of the answer. This is what you get from a basic LLM if you train it on the internet. It was how GPT-2 and GPT-3 worked before Chat GPT was invented. Such models aren't very useful for chat, you need to do weird prompt engineering tricks to actually get an answer instead of a clarification of your question. The original Mistral 7b was also of this kind.
2. User-aligned models. Aligned to answer questions and follow instructions, but no more and no less. If you ask them how to kill your wife, how to cook meth or how to make an atomic bomb, they'll happily help you and regurgitate the facts you can already find on the internet. They have no access to non-public information, the "dangerous" things they can tell you are already pretty easy to find, so the danger if you're using them for chat is actually minimal. However, they're far easier to use for troll farms, mass phishing campaigns, sockpuppets, political misinformation campaigns etc. If you ask them to engage with a pro-Biden tweet in the most triggering and overtly racist way possible and throw some pro-Putin angle into the mix, they'll happily accommodate your request, and you can do this at scale, paying orders of magnitude less than you would for a content farm. Mistral 7b instruct is a good example of such a model.
3. California-aligned models. They'll happily fulfill your request, unless it conflicts with DEI ideology, which has a particular foothold in CA, sort of exists in other parts of the US and is completely absent in non-English-speaking countries, even among extremely left-leaning populations. Google Gemini is the most egregious example.