Comment on GPT‑Red: Unlocking Self-Improvement for RobustnessComments−jing099281moUseful direction, but the hard part seems to be measuring novelty after each fix. Are they reporting whether later red-team cases are genuinely distinct, or mostly variants of the same failure mode?
Comments
Useful direction, but the hard part seems to be measuring novelty after each fix. Are they reporting whether later red-team cases are genuinely distinct, or mostly variants of the same failure mode?