When we notice a pattern we attempt to setup a question or prompt that will induce the pattern even though it is the incorrect solution, e.g., if we observe that the model tends to give long-winded answers we prompt the model to “Answer yes or no without explanation.” We use this to find “lemons” where we manage to adversarially break the model and “cherries” where we fail to break the model, and present both.
Comments
Is lemon-picked a real phrase or did they use GPT to generate the abstract? The term is “cherry-picked”.
I think the idea is they're showing worst examples (lemons) rather than best (cherries).