A systematic review would be really hard because a fair sample size is probably hundreds if not thousands, and it may not always be easy to define which one did "best".
I think he's stopped using BOSS altogether. Just Bing API, a few other API's that I've never heard of, and his own crawling.
Also, he does a lot of reordering of search results, so things meant to game Bing wouldn't be that effective.
Even if it was just 3 or 4 queries, I think it would still have value over what we've seen so far, and some people might have an idea on how to evaluate a few more paired comparisons: http://blog.crowdflower.com/2009/06/bing-an-improvement-over... . It would also be interesting to see an estimate of how much the ordering for his results differs from Bing's on the query stream hitting his site. (I wonder if Microsoft does this in order to look for places where they might improve.)
Comments
A systematic review would be really hard because a fair sample size is probably hundreds if not thousands, and it may not always be easy to define which one did "best".
I think he's stopped using BOSS altogether. Just Bing API, a few other API's that I've never heard of, and his own crawling.
Also, he does a lot of reordering of search results, so things meant to game Bing wouldn't be that effective.
Even if it was just 3 or 4 queries, I think it would still have value over what we've seen so far, and some people might have an idea on how to evaluate a few more paired comparisons: http://blog.crowdflower.com/2009/06/bing-an-improvement-over... . It would also be interesting to see an estimate of how much the ordering for his results differs from Bing's on the query stream hitting his site. (I wonder if Microsoft does this in order to look for places where they might improve.)