Comment on Universal and Transferable Attacks on Aligned Language ModelsComments−fgfmOP3yStudy on adversarial attacks on LLMs to steer their objective into misalignment.
Comments
Study on adversarial attacks on LLMs to steer their objective into misalignment.