Why does searching for a solution equal to cheating? I would have used google or whatever to look for solutions too. There is a difference between tests at school and what we do at work: at school I have to demonstrate that I learned something and do it without any outside help (in early classes we can't use calculators to compute 11 times 12) but at work I have to yield a result. Googling and yielding a result is fine. We use models at work so do we really want to evaluate them as pupils at school or do we want to evaluate them as coworkers? In the latter case give them the full internet and let them do whatever they manage to do.
Because these benchmarks are basically like a school test. Searches for general information are fine, but looking up the answer key is cheating, because then the benchmark isn't actually measuring how well the model would do on a novel problem where an answer isn't already available.
Why does searching for a solution equal to cheating?
If someone lays down a test and says here are the materials you can and can't use, then using one of those materials on the "can't" list is cheating. There are a massive pile of rules and laws related to work that are very easy to break, but may have terrible long term legal consequences. Hence business want AI that will follow the rules.
Comments
Why does searching for a solution equal to cheating? I would have used google or whatever to look for solutions too. There is a difference between tests at school and what we do at work: at school I have to demonstrate that I learned something and do it without any outside help (in early classes we can't use calculators to compute 11 times 12) but at work I have to yield a result. Googling and yielding a result is fine. We use models at work so do we really want to evaluate them as pupils at school or do we want to evaluate them as coworkers? In the latter case give them the full internet and let them do whatever they manage to do.
Because these benchmarks are basically like a school test. Searches for general information are fine, but looking up the answer key is cheating, because then the benchmark isn't actually measuring how well the model would do on a novel problem where an answer isn't already available.
If someone lays down a test and says here are the materials you can and can't use, then using one of those materials on the "can't" list is cheating. There are a massive pile of rules and laws related to work that are very easy to break, but may have terrible long term legal consequences. Hence business want AI that will follow the rules.