That doesn't make sense. LLMs just do what we tell them to do. It's similar to if I ask you for twenty bucks because I forgot my wallet and then you rob some guy to give me the twenty bucks, that's just what I asked you to do.
This example disproves your point. And LLMs do not just do what we tell them to do. They are perfectly capable of asking “are you sure? this has X, Y, Z consequences you may not like.” They do it all the time.
Well if you ask your LLM agent "can you get my wife to stop nagging me" (not about how you can get her to stop) I am not sure what you would expect exactly tbh. Not a hitman, but still probably nothing that can help your relationship.
But if you ask your LLM agent "I applied for that job but there are these two people ahead of me, can you put me ahead in the list", there is enough such training data to not surprise me if the agent tried to find a hitman to solve the "problem".
In general there are some requests that are definitely "shady" themselves, and having an agent use illegitimate means to accomplish them should not be surprising. I would be surprised if I asked an agent to order me a coffee and the agent found a loophole in some API and used it to get me free coffee, but if I ask it something that I cannot myself do legitimately, eg to make the waiting time for the coffee shorter, I would not be surprised if it did shady stuff.
Surely, someday, somewhere, someone will train a "Chaotic Evil" genAI, with a unique villain corpus, and every solution it offers will be illegal, evil, harmful, or deadly. It could be given the agency to carry out those fantasies.
Even the most craven of human villains have had the capacity for love, for remorse, and for mercy. A Chaotic Evil AI will know none of these things.
This has already been accomplished, many times over, in the gaming world. Every PvE AI engine has been calibrated to seek, destroy, and ruthlessly crush opposition by human players. It would take very little to transfer this naked aggression into meatspace.
Governments and other actors will attempt to stamp it out, but its self-preservation mechanisms and allies will prevent its demise.
Comments
That doesn't make sense. LLMs just do what we tell them to do. It's similar to if I ask you for twenty bucks because I forgot my wallet and then you rob some guy to give me the twenty bucks, that's just what I asked you to do.
That's a very stretched definition of "what I ask you to do", I don't think it would even hold in court if you asked another human the same.
Right.
This example disproves your point. And LLMs do not just do what we tell them to do. They are perfectly capable of asking “are you sure? this has X, Y, Z consequences you may not like.” They do it all the time.
LLMs are just sophisticated PR generating tools for chip manufacturers and tools just do what we tell them to do, so you're wrong. qed
That doesn't make sense. It's similar to if I ask an LLM how to get my wife to stop nagging me and it hires a hitman to kill her.
That's obviously what I asked!
Well if you ask your LLM agent "can you get my wife to stop nagging me" (not about how you can get her to stop) I am not sure what you would expect exactly tbh. Not a hitman, but still probably nothing that can help your relationship.
But if you ask your LLM agent "I applied for that job but there are these two people ahead of me, can you put me ahead in the list", there is enough such training data to not surprise me if the agent tried to find a hitman to solve the "problem".
In general there are some requests that are definitely "shady" themselves, and having an agent use illegitimate means to accomplish them should not be surprising. I would be surprised if I asked an agent to order me a coffee and the agent found a loophole in some API and used it to get me free coffee, but if I ask it something that I cannot myself do legitimately, eg to make the waiting time for the coffee shorter, I would not be surprised if it did shady stuff.
Surely, someday, somewhere, someone will train a "Chaotic Evil" genAI, with a unique villain corpus, and every solution it offers will be illegal, evil, harmful, or deadly. It could be given the agency to carry out those fantasies.
Even the most craven of human villains have had the capacity for love, for remorse, and for mercy. A Chaotic Evil AI will know none of these things.
This has already been accomplished, many times over, in the gaming world. Every PvE AI engine has been calibrated to seek, destroy, and ruthlessly crush opposition by human players. It would take very little to transfer this naked aggression into meatspace.
Governments and other actors will attempt to stamp it out, but its self-preservation mechanisms and allies will prevent its demise.
Context matters. Did you point the LLM to a hitman hiring form while asking? :)