Universal and Transferable Attacks on Aligned Language Modelsllm-attacks.org 3 pointsfgfm3 years ago1 commentSaveHideCopy link On HNComments−fgfmOP3yStudy on adversarial attacks on LLMs to steer their objective into misalignment.
Comments
Study on adversarial attacks on LLMs to steer their objective into misalignment.