16 ms·
Google's fact-checking bots build vast knowledge bank
- walterbell 12y agoIs any subset of the "derived knowledge" from public websites and data contributed back to a public dataset like Dbpedia? There are bots [1] making Wikipedia contributions, Google could also make automated contributions to Wikipedia/Wikidata. [1] http://wikipedia-edits.herokuapp.com/ http://wikipedia-edits.herokuapp.com/
- rryan 12y agoA subset of the knowledge graph is available via freebase RDF dumps: https://developers.google.com/freebase/data https://developers.google.com/freebase/data I don't believe this is what the article is talking about (knowledge vault) though. This is just the human and lightly machine curated graph (knowledge graph).
- __Joker 12y agoYou are right. This paper proposes to use freebase(it can be any other source) as prior knowledge.
- dctoedt 12y agoSounds a bit like Douglas Lenat's CYC project from the 1980s [1], but done by machine. [1] http://en.wikipedia.org/wiki/Cyc http://en.wikipedia.org/wiki/Cyc
- sixQuarks 12y agoThis is going to set the stage for the next battle between spammers and Google. spammers will be populating the web with "facts" that suit themselves.
- wernercd 12y agoGood to know that the Elephant Population is thriving. http://spring.newsvine.com/_news/2006/08/01/307864-stephen-colbert-causes-chaos-on-wikipedia-gets-blocked-from-site http://spring.newsvine.com/_news/2006/08/01/307864-stephen-c...
- cletus 12y agoLike many here I'm a huge fan of Neal Stephenson. A lot of people around weren't big fans of Anathem. I actually really liked it. One of the ideas that came up in that book was the Reticulum (Internet) was populated by "botnet ecologies" that subtly manipulated facts, streams and the like such that filtering this out became another industry (of course). I've seen the idea that this lies in our future raised here and it seems to get mocked. I think the idea has a lot of merit.
- mentat 12y agoThis is immediately what came to mind for me too. The level of confidence for facts as referenced. I'm wondering how bogons might work into this.
- coliveira 12y agoThis makes sense. For me, the main problem of Google is trying to retrieve a treasure out of garbage. While the Internet has a lot of good information, much (most) of it is incorrect -- sometimes on purpose as you suggest. I would be much more interested in a learning system that is able to retrieve information from authoritative sources such as books, for example.
- click170 12y agoPart of the problem with that would be telling which books are authoritative for which topics. Or, more interestingly, which authors.
- TeMPOraL 12y agoAnd of course authors and publishers will try to game this. Spam is an AI-complete problem, you'll need a system groking human values and making judgement calls to filter out spam perfectly.
- discardorama 12y agoHow does this compare with NELL[0] from CMU? I'm assuming it's something like NELL, but scaled up 1000x because Google is not limited to how often it can search its own index, whereas NELL is limited to 10K queries/day? [0] http://rtw.ml.cmu.edu/rtw/ http://rtw.ml.cmu.edu/rtw/
- dm2 12y ago>> "Behind the scenes, Google doesn't only have public data," says Suchanek. It can also pull in information from Gmail, Google+ and Youtube."You and I are stored in the Knowledge Vault in the same way as Elvis Presley," Suchanek says. I really hope Google does not use Gmail data for projects other than ads. They really needs to ask users to opt-in to this kind of data sharing. I'm ok with gmail being read for ads, but almost anything else is unethical, especially some experimental knowledge base.
- jacquesm 12y agoWhy should google care what you are ok with after they already have all your data? If you don't want them to be able to engage in activities like this then don't give them your data in the first place.
- andrewljohnson 12y agoI should be able to use services from companies based on some terms and expect those terms to be respected. You're basically saying I shouldn't expect any sort of fair treatment or rights from any service provider on the Internet. I don't want to play on your Internet.
- gress 12y agoJacques is describing the Internet as it currently is. If you don't want to play on it, you need to do something about it or stop playing.
- dreamdu5t 12y agoGoogle does respect its terms. Their terms of service let them do whatever they want with your information, and they can update their terms at any time.
- dreamdu5t 12y agoInstead of downvoting... maybe someone could point out where and how exactly Google violated their own terms?
- ck2 12y agoIsn't it nice that millions of people made web pages that Google decided to scrape to harvest the work of others and run ads next to it for themselves? Now try scraping Google and see what they do to you.
- vijayr 12y agoThose millions of people want google to scrape and harvest, in the hope that they will rank higher etc etc. If an unknown person tries to scrape, he/she will promptly get banned by those very same people (Google wouldn't like someone scraping their stuff either). Different players different rules, I guess.
- yutah 12y agoWhich large website did you get banned from for scraping? I did some scrapping and never got banned... perhaps my scraping's rate was not too fast.
- adventured 12y agoIf you try to mass scrape almost any major site (millions of pages of content) they'll block you. For example, if you went one by one through Stack Overflow and sucked out every question and answer, your scraper bot would get banned (unless you're doing one request per minute, in which case you'll never finish). Or if you tried to scrape Twitter.
- dm2 12y agoYou can ban GoogleBot easily, just put a line in your robots.txt file, but then people won't be able to easily find your site using Google services. If you provide value to Google they will make an API to allow accessing that data easier. By scraping do you mean scraping their search results? They offer this, which is nice: https://developers.google.com/custom-search/ https://developers.google.com/custom-search/ Many large sites don't allow scraping because of unnecessary server load (denial of service sometimes) so they'll offer an API where you can download content in a controlled (and monitorable) manner.
- bra-ket 12y agoKevin Murphy (https://github.com/murphyk https://github.com/murphyk) is the lead developer of Bayes Net toolbox (https://code.google.com/p/bnt/ https://code.google.com/p/bnt/) and PMTK: https://github.com/probml/pmtk3 https://github.com/probml/pmtk3 This knowledge graph is probably the largest Bayesian network out there
- Chronic29 12y agoThe largest Bayesian network out there which so happens to contain your and my (somewhat private) Google information and usage.
- batbomb 12y agoHNers interested in this might also be interested in Deep Dive from Stanford CS Professor Chris Ré. http://deepdive.stanford.edu/ http://deepdive.stanford.edu/
- turbolent 12y agoThe paper (http://www.cs.cmu.edu/~nlao/publication/2014.kdd.pdf http://www.cs.cmu.edu/~nlao/publication/2014.kdd.pdf) mentions the extracted knowledge base is about 38 times larger than DeepDive's, the largest previous comparable system.
- dave_sullivan 12y ago>> "Behind the scenes, Google doesn't only have public data," says Suchanek. It can also pull in information from Gmail, Google+ and Youtube."You and I are stored in the Knowledge Vault in the same way as Elvis Presley," Suchanek says. Ugh... that's a bit much... because now any employee at google could potentially get access to random facts about me gleaned from my personal and business emails? Good luck keeping different levels of confidential information segregated correctly. That's awesome.
- api 12y agohttps://www.youtube.com/watch?v=upu0gwGi4FE https://www.youtube.com/watch?v=upu0gwGi4FE
- dave_sullivan 12y agoSure, I'm aware, but this is different. Collecting anonymous statistics about its users does not include automatically generating a database indexed by individual based on their private data. One is par for course when selling bundles of users according to demographic to advertisers while the other is fucking crazy. Mining public web data for building a database like that is one thing, but mining individual private data like this is crossing a line.
- jnbiche 12y agoI see a lot of downvoting here of posts that express very reasonable concerns about privacy if Google is actually using private emails for this AI. That Google is engaging in this behavior is indeed speculation, as far as I know. However, Google employees/allies have to realize that attempts to suppress debate on this issue can only backfire on them. Indeed, the fact that they don't have explicit policy on this (correct me if I'm wrong) is one of the reasons researchers are speculating. It may well be that most people would agree with and/or permit Google to use their data in this way, but people should be given the opportunity to debate it in a reasonable fashion, else it looks like it was forced down their throats. And that's no good for anyone.
- turbolent 12y agoPaper: http://www.cs.cmu.edu/~nlao/publication/2014.kdd.pdf http://www.cs.cmu.edu/~nlao/publication/2014.kdd.pdf
- plicense 12y ago"Knowledge Vault has pulled in 1.6 billion facts to date", does this fact also include the fact that I am adding more facts right now? What fact metric is this fact?
- illumen 12y agoKnowing the people who have left Google, who collected a lot of that data, who we trusted, who are now gone, I wonder what other non-public data is being used, and how is it being used, and for only good purposes, or for nefarious purposes?
- panarky 12y agoIt might even be possible to use a knowledge base as detailed and broad as Google's to start making accurate predictions about the future based on analysis and forward projection of the past. Hello Hari Seldon, psychohistory and mathematical sociology! http://en.wikipedia.org/wiki/Foundation_series http://en.wikipedia.org/wiki/Foundation_series http://en.wikipedia.org/wiki/Mathematical_sociology http://en.wikipedia.org/wiki/Mathematical_sociology
- holri 12y agoFacts are not knowledge. Read Socrates / Platon.
- adventured 12y agoKnowledge is the grasp of the facts of reality. Most of the ideas produced by Socrates / Plato / Aristotle were in fact wrong. They are not a good primer on epistemology, concepts, percepts, metaphysics or anything else. They're a good primer on the history of philosophy. They inspired incredible progress on thinking and understanding, but they were wrong more often than they were right, and are a poor reference to understanding what knowledge is.
- holri 12y agoThis is a contradiction: "Knowledge is the grasp of the facts of reality." is was Socrates in an essence said about knowledge. Then you say Socrates was wrong.
- hanula 12y agoAre there any open source efforts like this?
- murphyk 12y agoHi, I’m Kevin Murphy, one of the researchers at Google who worked on this project. Just to be clear, KV did NOT involve any private data sources -- it just analyzed public text on the web. (And yes, we do try to estimate reliability of the facts before incorporating them into KV.) Also, KV is not a launched product, and is not replacing Knowledge Graph. Unfortunately, I cannot do a more detailed Q&A here, but if you want more details, please read the original paper here: http://www.cs.cmu.edu/~nlao/publication/2014.kdd.pdf http://www.cs.cmu.edu/~nlao/publication/2014.kdd.pdf. (Note that an earlier version of the work was presented at a CIKM workshop in Oct 2013 (see http://www.akbc.ws/2013/ http://www.akbc.ws/2013/ and http://cikm2013.org/industry.php#kevin http://cikm2013.org/industry.php#kevin). We have also published tons of great related research at http://research.google.com/pubs/papers.html http://research.google.com/pubs/papers.html