Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Excellent analysis, thank you, especially in the distinction between "suggested sites" and the Bing toolbar's behavior. I can't argue with your methods. However, I do differ with some of your conclusions:

> The behaviour I’ve seen explains Google’s experiments, but does not support the accusation that Bing set out to copy Google.

I don't think it's so much about "set out to copy Google" necessarily, as it is that they are explicitly parsing Google queries and results from the clickthrough data and using it (quite directly) for their own results. What they set out to do is immaterial given what is provably happening.

> Bing Toolbar is tracking user clicks and Bing could use the result to improve search results. I don’t personally see any great distinction between this behaviour and Google’s many tracking, indexing and scraping endeavours which they use to improve their own search results.

The difference is that Google has proven that the results of certain queries are being directly fed as Bing results. If Microsoft does the same with Google rankings, I'd see your point, but right now the evidence only points in one direction.

> While I personally dislike the privacy implications, Bing Toolbar is pretty upfront about it when it gets installed (unlike much web page user tracking.) The fact that the tracking is plain HTTP not HTTPS, with the content in plaintext, would seem to indicate that they weren’t seeking to hide anything.

I'd be interested to see if Google over HTTPS queries are being transmitted by the toolbar over HTTP. That would be a pretty serious privacy violation IMO, especially when you pair that with unencrypted wifi at Starbucks. See the AOL search log fiasco: http://www.somethingawful.com/d/weekend-web/aol-search-log.p...



The difference is that Google has proven that the results of certain queries are being directly fed as Bing results. If Microsoft does the same with Google rankings, I'd see your point, but right now the evidence only points in one direction.

However, these are also the only experiments that have been run - outlier data, gaming the algorithms. If you only test one possible outlier scenario and don't control against any other, it's fairly ambitious to stand up and say "this is exactly what is happening!"

IMHO it would have been better for Microsoft to respond along these lines, instead of going into counterspin mode, though.


Claiming that Bing is copying outlier results is still a fair claim that Bing is copying results, especially considering that outlier queries are hard in general for search engines to get right: http://portal.acm.org/citation.cfm?id=1277939


what control are you suggesting? I read a suggestion elsewhere of using only Firefox for some queries and seeing if they show up, which is silly, of couse.

without a possible mechanism of action there's no point of using something as a control; I might as well wait to see if the files on my disabled usb drive show up on bing. they were "gaming the algorithms" because they suspected that the algorithms existed, and it wouldn't have worked if there was nothing to game.


Set up some other dynamic sites (not indexed or linked anywhere) with random nonsense words in their URL strings, and a single link to some unrelated site. Click the link. Repeat, wait two weeks.

Same as what google did, just not on google.com.

Of course, if Bing has other "relevance" indicators (ala BingPageRank) then this might not work because it will rule the nonsense control site irrelevant linkspam, and not put it in (even Google's experiment only got a 9% success rate.) It would be better if someone like Facebook injected the test nonsense strings in their URLs.

Something of this sort would at least show they tried to differentiate between "scrapes all clickstream data" and "uses Google's search results".




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: