It gives me a little pause that humans are so much worse than random chance at detecting GPT-4.5. Suppose we reframed the test as: "You interact with 10 witnesses, 5 of which are humans, 5 of which are GPT-4.5. Your task is to separate them into two groups, but you do not need to label the groups." It seems that human judges would still be pretty good at this version of the task.
In originally proposing the task, Turing wrote:
It might be urged that when playing the "imitation game" the best strategy for the machine may possibly be something other than imitation of the behaviour of a man. This may be, but I think it is unlikely that there is any great effect of this kind.
Does the fact that GPT-4.5 is favored well above random chance imply that it is doing "something other than imitation of the behaviour of a man"?
Comments
It gives me a little pause that humans are so much worse than random chance at detecting GPT-4.5. Suppose we reframed the test as: "You interact with 10 witnesses, 5 of which are humans, 5 of which are GPT-4.5. Your task is to separate them into two groups, but you do not need to label the groups." It seems that human judges would still be pretty good at this version of the task.
In originally proposing the task, Turing wrote:
Does the fact that GPT-4.5 is favored well above random chance imply that it is doing "something other than imitation of the behaviour of a man"?