And while in general that is an incredibly difficult and complex problem, for most benchmark cheating it seems almost trivial: run the benchmark in a vm that has neither network access nor access to the scoring code. For remote models use a proxy that proxies exactly that one endpoint to call the llm, and rejects any calls that configure provider-side tooling (since e.g. OpenAI has their own WebSearch you have to prevent the model from using)
Comments
And while in general that is an incredibly difficult and complex problem, for most benchmark cheating it seems almost trivial: run the benchmark in a vm that has neither network access nor access to the scoring code. For remote models use a proxy that proxies exactly that one endpoint to call the llm, and rejects any calls that configure provider-side tooling (since e.g. OpenAI has their own WebSearch you have to prevent the model from using)