The reason is that with the URL and the query, there's a decent chance you can derive the association.
Now the reason I say this is that to my best approximation the key difference that makes Google better than Bing (when it is) is the index size, and not the search algorithm.
So the two key pieces of value are:
A) queries that people don't do on your site, which would have generated poor relevance
B) Pages that are actually browsed that you don't have indexed or up to date.
Now don't get me wrong. There is value in the association, but I think I'd capture 80% of the value with the above. Now as you point out, these lists are often non-obvious, but Bing does a near equal job of creating the lists. And if you give them the search terms where they have gaps. And fill the index, then I think the gap closes most of the way.
And lets be clear, because they clicked the link doesn't mean its a good link. See ehow.com or expertsexchange. But it does have some value.
And frankly I think MS would have been willing to let it go if Google had gone straight to MS with it and said, "We have this data. Even though its not unethical we can spin it to the media to make it look bad. Kill the association." MS probably would have as the net win isn't that huge. Now the PR loss in pulling it would be worse than the PR loss in keeping it (with no integrity loss, since they [and I] think it is perfectly ethical).
I'm pretty thoroughly sure this is false. I'm sorry I can't show data to back it up, but Google has many many systems that are years ahead of bing's technology. I'm kinda hamstrung here by being unable to reveal anything about them. To a decent approximation, both Google and Bing likely have the whole internet that matters (and that they are allowed to crawl) in their indexes. Bing is just unable to return this data for as many queries as Google is.
>but Bing does a near equal job of creating the lists.
To the extent that this is true, how much of it is due to data harvested from Google search? This is something that google can't really demonstrate with evidence, and what I would have hoped that bing would clarify in a public statement if they had anything defensible to say.
To a decent approximation, both Google and Bing likely have the whole internet that matters (and that they are allowed to crawl) in their indexes. Bing is just unable to return this data for as many queries as Google is.
Maybe this is true, but its not apparent. I had commented on this a month or so back that I thought Bing was better at "vague" queries, where I don't know exactly what I'm looking for. But I'll know it when I see it.
Whereas Google is really good at targeted queries. I need info on the HP battery model number 003D434F90. These searches in Bing will often bring back literally nothing, while Google will often have one or two links, but they happen to be the link that I want. The text of the query is almost always found in these pages.
From that I infer it is index size, since the text is in the page.
To the extent that this is true, how much of it is due to data harvested from Google search?
While I find the quality of searches similar I don't find the results to be similar, if that makes sense. If Bing is harvesting your results, they're still using other very clever methods to surface other equally good, yet different results.
One obvious and very public example of how the algorithm matters more than the index size would be Cuil. It was launched with lots of hype about how it would have a index several times larger than Google's. And of course the results were notoriously bad.
I'm stunned that anyone could think that the core issue in this controversy is indexing or page discovery. It should have been totally obvious from the examples in the initial article and blog post that this was specifically about copying ranking.
Now the reason I say this is that to my best approximation the key difference that makes Google better than Bing (when it is) is the index size, and not the search algorithm.
So the two key pieces of value are: A) queries that people don't do on your site, which would have generated poor relevance
B) Pages that are actually browsed that you don't have indexed or up to date.
Now don't get me wrong. There is value in the association, but I think I'd capture 80% of the value with the above. Now as you point out, these lists are often non-obvious, but Bing does a near equal job of creating the lists. And if you give them the search terms where they have gaps. And fill the index, then I think the gap closes most of the way.
And lets be clear, because they clicked the link doesn't mean its a good link. See ehow.com or expertsexchange. But it does have some value.
And frankly I think MS would have been willing to let it go if Google had gone straight to MS with it and said, "We have this data. Even though its not unethical we can spin it to the media to make it look bad. Kill the association." MS probably would have as the net win isn't that huge. Now the PR loss in pulling it would be worse than the PR loss in keeping it (with no integrity loss, since they [and I] think it is perfectly ethical).