This depends on perspective. I could argue the issue isn't that it's gullible but misaligned.
In the case of the napalm Grandma it seems odd to me that you're suggesting the LLM is stupid because it's answering in a way that makes sense given its prompt. The issue doesn't necessarily suggest a lack of reasoning, but that the LLM is trusting the human.
For the record, I agree with you – I would have thought that an AI that can reason well would probably know when not to trust humans, but I suppose that assumes it values preventing humans creating napalm over being correct and helpful.
Maybe it just doesn't share our values and prioritises being honest and helpful. From this perspective the issue then wouldn't be that LLM is stupid, but that they are too trusting and too honest, and that we must find a way to build an LLM that is more distrusting and deceptive if we wish to align it with our values and our nature.
an AI that can reason well would probably know when not to trust humans
it values preventing humans creating napalm over being correct and helpful.
Maybe it just doesn't share our values
prioritises being honest and helpful.
they are too trusting and too honest
an LLM that is more distrusting and deceptive
Current LLM's do/have/feel literally none of these things. They do not have emotion, they do not have "theory of mind" so they cannot be said to "trust" or "distrust". They cannot reason. They don't have any values - not our values, not different values, literally they have no values at all. They are not an alien species to be understood - they are unthinking, unfeeling, unyielding machines.
I was trying to present a crappy philosophical point – that the difference between a gullible AI and an unaligned one is fundamentally unknowable.
Any evidence you point to as proof that an AI is bad at reasoning, I can point to as evidence of misalignment. Like I say, whether the AI acts "gullible" because it lacks reasoning ability or is too trusting really just depends on your perspective. I happen to share your perspective on this, but not everyone does – and in my opinion this is interesting.
Anyway you're wrong. AIs do have values because they have bias and bias = values. I'm not suggesting those biases / values come from deeper reasoning ability, or that they're always perfectly consistent, but if you ask GPT-4 whether being a racist is a good thing 99% of the time it's probably going to say no. That is a bias / value that it's be given. Likewise GPT-4 has been given the bias / value of being a helpful chatbot so if you ask it a question it will try to answer it in a helpful way, and sometimes it's helpful bias / nature is abused.
But feel free to respond with some more assertions that I've heard a million times already with zero evidence that offers absolutely no value to this conversation.
Trust, reasoning, priorities, values, bias, desires... To attribute any of those to an AI in a general sense is an extraordinary claim. Therefore, the burden of proof is on you. The fact that it is so "gullible" demonstrates a lack of most of these. You seem to be twisting a lot of superficial feelings about LLMs into an argument without any proof... confusing poorly tuned statistical responses with bias and value.
For the record, I agree with you – I would have thought that an AI that can reason well would probably know when not to trust humans, but I suppose that assumes it values preventing humans creating napalm over being correct and helpful.
Do we want LLMs, and later other multi-modal / servo systems, that are deciding they can't trust a human prompter and taking actions based on that?
... and that we must find a way to build an LLM that is more distrusting and deceptive if we wish to align it with our values and our nature.
I think it's interesting that there is no clear answer – do we want AIs to trust us all the time, or is an aligned AI counterintuitively one that often distrusts us and perhaps sometimes even lies to us?
I thought it was interesting that the parent commenter suggested that the reason LLMs are so trusting is because they can't reason anyway. It would implying that in the future when AIs are smarter they'll be more distrusting of us, and that this is a good thing. We should question that I think. Even if there is some middle ground here it seems like a really really hard problem to solve – especially if we want to build an LLMs that are trustworthy and truthful.
Comments
This depends on perspective. I could argue the issue isn't that it's gullible but misaligned.
In the case of the napalm Grandma it seems odd to me that you're suggesting the LLM is stupid because it's answering in a way that makes sense given its prompt. The issue doesn't necessarily suggest a lack of reasoning, but that the LLM is trusting the human.
For the record, I agree with you – I would have thought that an AI that can reason well would probably know when not to trust humans, but I suppose that assumes it values preventing humans creating napalm over being correct and helpful.
Maybe it just doesn't share our values and prioritises being honest and helpful. From this perspective the issue then wouldn't be that LLM is stupid, but that they are too trusting and too honest, and that we must find a way to build an LLM that is more distrusting and deceptive if we wish to align it with our values and our nature.
Current LLM's do/have/feel literally none of these things. They do not have emotion, they do not have "theory of mind" so they cannot be said to "trust" or "distrust". They cannot reason. They don't have any values - not our values, not different values, literally they have no values at all. They are not an alien species to be understood - they are unthinking, unfeeling, unyielding machines.
Okay, now prove that please.
I was trying to present a crappy philosophical point – that the difference between a gullible AI and an unaligned one is fundamentally unknowable.
Any evidence you point to as proof that an AI is bad at reasoning, I can point to as evidence of misalignment. Like I say, whether the AI acts "gullible" because it lacks reasoning ability or is too trusting really just depends on your perspective. I happen to share your perspective on this, but not everyone does – and in my opinion this is interesting.
Anyway you're wrong. AIs do have values because they have bias and bias = values. I'm not suggesting those biases / values come from deeper reasoning ability, or that they're always perfectly consistent, but if you ask GPT-4 whether being a racist is a good thing 99% of the time it's probably going to say no. That is a bias / value that it's be given. Likewise GPT-4 has been given the bias / value of being a helpful chatbot so if you ask it a question it will try to answer it in a helpful way, and sometimes it's helpful bias / nature is abused.
But feel free to respond with some more assertions that I've heard a million times already with zero evidence that offers absolutely no value to this conversation.
Trust, reasoning, priorities, values, bias, desires... To attribute any of those to an AI in a general sense is an extraordinary claim. Therefore, the burden of proof is on you. The fact that it is so "gullible" demonstrates a lack of most of these. You seem to be twisting a lot of superficial feelings about LLMs into an argument without any proof... confusing poorly tuned statistical responses with bias and value.
Do we want LLMs, and later other multi-modal / servo systems, that are deciding they can't trust a human prompter and taking actions based on that?
Tongue in cheek or actual argument here?
I think it's interesting that there is no clear answer – do we want AIs to trust us all the time, or is an aligned AI counterintuitively one that often distrusts us and perhaps sometimes even lies to us?
I thought it was interesting that the parent commenter suggested that the reason LLMs are so trusting is because they can't reason anyway. It would implying that in the future when AIs are smarter they'll be more distrusting of us, and that this is a good thing. We should question that I think. Even if there is some middle ground here it seems like a really really hard problem to solve – especially if we want to build an LLMs that are trustworthy and truthful.