Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

As someone who has done research in the field, I can tell you - this is the future of search, and the key to true AI. Once we have a knowledge graph of the world, you will be able to ask questions and get answers, not documents. This will allow so much more automation - personal assistants like Siri will be able to "book me a room in San Francisco with a view of the ocean" or "give me a list of the schools all the children of American presidents went to" and so on.

There is an arms war here between Bing and Google - we can thank Microsoft for pressuring Google into making this happen sooner than they otherwise would have.



Isn't this exactly what people thought in the 80s? Building knowledge into a computer of the relationships between entities in the world will let you write expert systems that work like humans?


I read at some point that a lot of interesting theoretical AI research was abandoned in the 80s, because the results (like expert systems) got good enough that they had commercial applications, and the focus changed from research to perfecting existing products. No idea if that's true, but it suggests that there might be room to pick up where people in the 80s left off. Especially given that computing resources and data archives are somewhat improved since then.


Research funding was also cut heavily in late 80s: http://en.wikipedia.org/wiki/AI_winter


Might be true, but it sounds unlikely for two reasons:

(1) If new theories of AI proved to be commercially viable, you would expect that people would continue to develop new theories in the pursuit of money and fame, not "abandon" them because the results are "good enough."

(2) There are still lots of smart people working on AI in academia today that were around in the 80s. (And the 70s. And the 60s.) They remain keenly aware of the breakthroughs that occurred then, because they were the ones making them. And they're familiar with the technology of today, because they're using it in their new research. So if any big rocks were left unturned prematurely, I'm sure they're being examined.


I think the idea is, it's like the difference between searching for new cancer treatments and bringing a promising treatment to market. If it takes the same kind of expertise to do both, then spending more time on one means less on the other.

Still no idea if that's true of AI, but I found a couple of interesting cites:

Patrick Winston, director of MIT's Artificial Intelligence Laboratory from 1972 to 1997, echoed Minsky. "Many people would protest the view that there's been no progress, but I don't think anyone would protest that there could have been more progress in the past 20 years. What went wrong went wrong in the '80s."

Winston blamed the stagnation in part on the decline in funding after the end of the Cold War and on early attempts to commercialize AI. But the biggest culprit, he said, was the "mechanistic balkanization" of the field, with research focusing on ever-narrower specialties such as neural networks or genetic algorithms. "When you dedicate your conferences to mechanisms, there's a tendency to not work on fundamental problems, but rather [just] those problems that the mechanisms can deal with," said Winston.

http://www.technologyreview.com/computing/37525/

They don't go into detail about how the early attempts to commercialize contributed to the problem. This site does, but seems less trustworthy:

In the early 1980s, dark clouds also settled over the MIT Artificial Intelligence Lab as it split into factions by initial attempts to commercialize Artificial Intelligence (AI). In fact, some of MIT's best White Hats left the AI Lab for high-paying jobs at start-up companies.

http://computer.yourdictionary.com/golden-age-era

So it sounds like some smart people in academia in the 80s think that some stones were left unturned, or turned too slowly, and that part of the problem was a refocus on making money on existing discoveries. According to that AI winter link, the tech mostly wasn't ready for primetime yet, presumably making it even harder to raise funds for new research.


Winston is still teaching and doing research at MIT. In fact I took two of his classes a couple years ago, and he's exactly who I had in mind when I mentioned researchers with experience and knowledge from decades past continuing to work :) Even if some previous research wasn't fully fleshed out, we can be confident that it hasn't been forgotten.


Yes, and I - for one - am not convinced that that approach has been proven wrong so much as it just hasn't been realized yet, for whatever reason. If I had to guess, I'd say that "whatever reason" is some combination of not having sufficient hardware, and not fully understanding knowledge representation and reasoning algorithms well enough yet.

Of course I could be wrong, but I think there's a good chance that a lot of "Good Old AI" stuff is still valid, but that it was just too early for it then. Maybe it still is, time will tell.


It's not wrong. Past failure is not guarantee of future failure.

Elon Musk and the electric car industry is one example.


"past failure ... the electric car industry is one example."

Actually, electric cars were popular before gas cars were.

http://en.wikipedia.org/wiki/History_of_the_electric_vehicle...


Interesting thanks!

But I think your argument falls down, when we inspect the merit of the world "popular".

Electric cars, have never been mainstream.


Yes, but the "web of knowledge" folk were planning (ish) hand-added metadata to allow knowledge extraction by AI programs, with no real incentives or network effect to promote people doing that. Google got microdata (which are semantic annotation, if a lot less complicated than the old schemes) added to websites by saying they would include them in Pagerank calculation.


When did google say that they would factor microdata into SERP ranking?


After doing some hunting, apparently I imagined them saying that. What I did not imagine was that it impacts score in results for their "recipe search" product - and it caused some consternation among the recipe blogging community at the time.


The Semantic Web is truly the all of those things -- but this is not that. My guess is that this will just end up as a sharpened version of the data section of Wikipedia results that we have now.

Why? Because they are selling it without mentioning why they won't fail where every other attempt has. There are huge difficulties in this. Have they turned a corner on the research that changes something? I don't see it in what they've so hinted at so far.


I think that there are two things in their favor here:

First, they are Google, and therefore possess huge quantities of data and the ability, courtesy of their uber map reduce prowess and ultra-fast custom hardware, to make sense of it.

Second, they bought Metaweb (makers of Freebase) and with it some of the best semantic expertise out there. Toby Segaran is a brilliant dude. His O'Reilly book "Programming the Semantic Web" explains in 20 pages what most books take 150 pages to do: the concept of a URI based graph database and how it enables data to be merged from multiple sources and reasoned over with applications.

I only hope Google open-sources some of their research here for the rest of us.


Thanks so much! Glad you enjoyed the book. I wanted to point out that Colin Evans and Jamie Taylor were also authors (and still work on the Knowledge Team at Google) and should get some credit.


Haha! Thanks for pointing that out. I didn't mean to leave them off. I met you guys in 2009 at a semantic tech conference, back before Metaweb was purchased. So glad to see your work being pushed to the most popular website in the world.


There's little need. It's a direct application of concepts that are well treated academically in all the various datalog papers.


Taking academic research to production ready code is far from trivial.


I think that generalization is false as often as it is true.

But also, have you read any of the papers involved? Datalog is pretty simple. It's a restricted, forward chaining prolog. Once you know that, you can recreate most of it from that description alone.


Google is sitting on an enormous pile of semantic data that dwarfs anything that any AI project could possibly use 20 or even 10 years ago. Of course it does not guarantee that they won't fail as well, but it sure gives them an edge.


It also appears that SEO is finally the thing that gets people to actually publish semantically-marked data, so they'll get some of this more easily in the near future.

http://bergie.iki.fi/blog/google-s_rich_snippets_will_lead_u...


Yes, like the HTML web this web of data grows by network effects. The "Semantic Web" was only ever a data model and exchange formats - a standard to guide independent efforts. Now, of course, a lot of that data will be protected by the internet giants for their competitive reasons, but the standards still provide the interoperation at the edges.


This is Google's search team we're talking about here. They might make the mistake of overestimating how much people will like something like this. They might make the mistake of underestimating how many resources something like this will take. However, they simply don't make mistakes of the kind you're talking about. I'd be willing to bet money that there's a bullet-proof theoretical model behind what they're doing. And they probably aren't ready to talk about it yet.


You are falling for the fallacy of infallibility. You assume they are successful because they have some magic secret, not just a lot of hard work.


Except that I named two areas where they can make mistakes. And Google's success is indeed because of the hard work they've put into building such a smart team of search engineers.


I agree. This looks to me to just be Wikipedia presented next to the search term. Much the same as DuckDuckGo http://duckduckgo.com/?q=taj+mahal Congrats Gabriel, you got to them :)

I thought that while flawed cpedia (from one of the cuil founders) was a much more interesting push on this idea then Googles one currently is.


while I share your enthusiasm for the bright future ahead, there is a missing link between this and "book me a room in San Francisco with a view of the ocean", namely that data is behind closed doors.

There is a bunch of data that google can use[1] because it is made explicitly available. But many sources don't want that.

As an example, consider "book me flights for the cheapest route between lisbon and kiev". It is a trivial thing to do, provided you can get airline data.

But you can't scrape ryanair's website because they willingly put counter measures in place (e.g. captchas) so you cant do that.

[1] e.g. http://richard.cyganiak.de/2007/10/lod/


Could the google bot take a captcha it finds and use it in one of it's own re-captchas? Essentially passing the burden of deciphering the text onto some unsuspecting human, it will be able to beat all captchas with ease!


My stack level got too deep reading your comment.


"I'm on it, riffraff." "I've found the cheapest flight between lisbon and kiev. Booking is possible thru ryanair.com . I've filled in all necessary details for you, however there is a captcha I can't wrap my cpu around. Also, there are some privacy agreements I'm not authorized to make for you. Could you take over from here?"

Imagine an integrated Siri with those kind of capabilities. It doesn't have to be fully automatic. Letting a secretary do stuff also isn't fully automatic, (s)he's there to optimize your time into doing only the important decisions (sign on agreements, clicking confirm after having seen the price..).


Once bots get smart enough there will be no captchas that can stop them. The only way will be to ban the IP where the bot is coming from, assuming you know which IP belongs to which bots. They could easily disguise themselves by using proxies.

I guess my point is that if we get to the point were a bot can do generic requests without the aid of a human a captchas will probably not be able to stop it.


You're assuming Google would do that and/or that it will be allowed to do that once sites find out. They are a lot of businesses who have zero incentive to give their data to Google.

This "AI" is just scrapping and replacing wikipedia while serving ads.


I think Google bought a company related to this technology 2-3 years ago. It's not like they saw Bing's Facebook integration last week and then decided to do this.



I'm not sure if I would want to say "a knowledge graph of the world" is the "key to true AI".

I think the situation is closer to "true AI is the key to a usable if situationally depend knowledge graph". Because the world doesn't have single knowledge graph that you can learn and use in all situations. Certainly, you can find a lot of common instances where the average works but once you're past that, you need the kind of understanding of language that present day systems are far from having.


"Because the world doesn't have single knowledge graph that you can learn and use in all situations."

The logical way to overcome this via a "data first" brute force approach is to build personalized knowledge graphs of every potential customer. Which is in effect what every statistically sophisticated large business is attempting.


Just to be nitpicky, don't you mean "the key to true", expert-system based AI? Or, are you going the "human intelligence is a clustering system; therefore, true AI would have a similar method" route?

Both are valid, I'm just curious.


Bing? I Use it not much, have they done something in that direction?

I like WolframAlpha.


They bought powerset a few years back and they did that.


>As someone who has done research in the field, I can tell you - this is the future of search, and the key to true AI.

Having done research in this field doesn't qualify you to have authoritative opinions on search or AI. I don't necessarily disagree with you on search, but I don't see why a dataset would be the key to AI. AI is a function, not a dataset.


Actually Cyc is panned pretty hard as a complete waste of a perfectly intelligent researcher's time.


Unless the hotel booking system has a captcha...

I always try to mention his when people get over excited a out the semantic web. Most of the important stuff that we would like to automate we also insist that a bot not be allowed to do it. It's pretty scitzo in my opinion.


Presumably the booking would be done though an API.


I'd say thank Metaweb as it seems to be building directly upon their ideas. I'm not sure how much Bing has pressured Google to innovate in the search space (though I think Bing has got Google thinking more about design).


I hope you enjoy your very pleasant, if initially surprising, stay in San Francisco, Cebu. Watch out for sea snakes!

http://en.wikipedia.org/wiki/San_Francisco,_Cebu

As someone who has done research in the field, did you read the Gizmodo review of Siri?

http://gizmodo.com/5864293/siri-is-apples-broken-promise

The set of "knowledge graph" problems which is not, in fact, AI-complete, strikes me as much smaller than most doing research in the field would like us to think.

I hope you won't try to argue that "book me a room" etc. isn't AI-complete. There's an uncanny valley there, and it's deep. You can use Siri for lots of trivial tasks in which bizarre failures are hilarious rather than disastrous, but booking hotel rooms isn't in that set. She can be the best secretary in the world 9 times out of 10 or even 99 out of a 100, but the other times she's an insane robot who wouldn't at all mind sending you to Cebu to get kidnapped by the MILF...


These are still early days. Siri is slow and inaccurate. However consider it a preview of things to come.


These are certainly the early days. SHRDLU was slow and inaccurate. In many way, it was an even impressive example of "things to come" (it could seem to understand some fairly complex figures of speech for example). The problem is that we're not sure example when these "things to come" will actually appear.

http://en.wikipedia.org/wiki/SHRDLU




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: