I think it's interesting that there is no clear answer – do we want AIs to trust us all the time, or is an aligned AI counterintuitively one that often distrusts us and perhaps sometimes even lies to us?
I thought it was interesting that the parent commenter suggested that the reason LLMs are so trusting is because they can't reason anyway. It would implying that in the future when AIs are smarter they'll be more distrusting of us, and that this is a good thing. We should question that I think. Even if there is some middle ground here it seems like a really really hard problem to solve – especially if we want to build an LLMs that are trustworthy and truthful.
Comments
I think it's interesting that there is no clear answer – do we want AIs to trust us all the time, or is an aligned AI counterintuitively one that often distrusts us and perhaps sometimes even lies to us?
I thought it was interesting that the parent commenter suggested that the reason LLMs are so trusting is because they can't reason anyway. It would implying that in the future when AIs are smarter they'll be more distrusting of us, and that this is a good thing. We should question that I think. Even if there is some middle ground here it seems like a really really hard problem to solve – especially if we want to build an LLMs that are trustworthy and truthful.