GPT‑Red: Unlocking Self-Improvement for Robustnessopenai.com 35 pointsalvis1 month ago1 commentSaveHideCopy link On HNComments−jing099281moUseful direction, but the hard part seems to be measuring novelty after each fix. Are they reporting whether later red-team cases are genuinely distinct, or mostly variants of the same failure mode?
Comments
Useful direction, but the hard part seems to be measuring novelty after each fix. Are they reporting whether later red-team cases are genuinely distinct, or mostly variants of the same failure mode?