Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

>You can't claim ownership of facts, even if you discovered them first, or created them into existence.

It's not as straightforward as all that. Suppose you write a book. The existence of that book is a fact. The words that are in that book are now a fact. Now, I create a book that says, "The following book was written by DennisM, this is a list of the words that are in that book, in order" and then precede to duplicate your book. Those are the words that are in the book, it's a fact, no one can deny that it's a fact. That's not defensible, even though I'm just stating a fact.

We're in a somewhat grey area here, dealing with ethical issues that have never been clearly dealt with before. However, I feel that a general principle of the web applies: If I create some data, and don't at least implicitly give you permission to use it, it's unethical for you to use it.

This is the central idea behind copyright, behind robots.txt, behind plagiarism, and even behind privacy.



In your analogy there was a wholesale copyright infringement, entirely absent from bing sting. Here's a much closer analogy:

I wrote a book, in which I described some new way to dance. Before the book was published, the ideas were mine to keep secret. After the book was published, a bunch of readers picked up on the idea. They started all dancing in certain way, and someone described their behavior. That description will be de-facto copy of my book, and yet it is entirely legitimate. Because users own they behavior, not the author of the book who inspired them.

Now if someone simply copied pages from my book, that would be a copyright violation. But that's not what happened. As it is, it's an original work of art. Would that suck for me as a dance-inventor? Obviously. Do I want to live in a society where description of my behavior is owned by the person who inspired or directed it? Absolutely not.

Your idea of data ownership is contrary to tradition and contrary to the law. Facts (such as relationship between FOO and BAR) can not be owned in any way shape or form, except as a trade secret. You can not copyright a fact, or trademark a fact. In some cases you can patent application of a fact to a problem, but that's not the case here, as you can't and don't want to patent relationship between a word and target page.

Just because you put effort into something, does not mean it's yours. You probably wish it were true, but again, that's not how the law works, and not what the tradition is. You can't own facts.


Actually, it is as straightforward as all that. What's alleged is that someone searches for DenisM's book on Google, and when they click on one of the results, just that piece of data is forwarded to Microsoft. Not a scraping of all of Google's search results and nothing like the contents of the book.

As to "the central idea of copyright", bear in mind that in addition to not being able to copyright ideas or data, there's also the notion of "fair use".


>nothing like the contents of the book.

Explain this to me. In terms of raw bytes it's probably more data than every book that has ever been written. It was certainly more expensive to write than any particular book. What is the important difference here?

To make it a little more concrete, suppose someone was doing the same thing with google's streetview data. Now, it's an empirical fact that if you go to such-and-such an address, and look around, it looks like the pictures that google took. Does that mean it's ok to use those pictures, just because their resemblance to reality is factual?


In terms of raw bytes it's probably more data than every book

You're referring to the aggregate, so far as I know copyright law works on specific instances. Let's say the Kindle had a button that allowed you to isolate a couple of sentences and broadcast them to your friends on Facebook, or wherever. The small portion of the text being shared would fall under fair usage and not be a copyright infringement violation, even though the aggregate of all Kindle users might be large.

What is the important difference here?

I think the biggest difference between the two sides debating this issue is that one side thinks the amount of effort Google put into creating search results is a relevant concern, while my side thinks that once those results are shared with users Google's legal/moral claims to any ownership of that data evaporate.

The streetview point highlights the problem. If Google and I both take pictures of the Church of Mxyzptlk at 14th and Main, we both own copyright to our respective pictures. The fact that I created my picture after "piggybacking" off of a Google search for Mxyzptlk is irrelevant to copyright. So far as know (IANAL) the combination of a link with search terms is a utilitarian set of data rather than a creative expression and hence not copyrightable.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: