10 ms·
Show HN: A database of everything (over 55M keys)
- ethernetsalad 10y agoAs much as I like the idea, I can already see someone called "Drumpf" on the front page linking attributes about Hillary Clinton and Donald Trump to their Twitter accounts and not their names. I guess if you're after "everything" then curation makes no sense but you'll end up with a bunch of nonsense.
- outpan 10y agoTwitter url seems like a strong key since it is referenced in many other contexts. As for curation, it is intended to provide examples of what key, attr and values are regardless of their content. This is an experimental feature and might be removed/tweaked...
- tisryno 10y agoGreat concept but I can see it falling out of line very quickly, on the homepage I spotted "berlin -> country -> germany" Followed by "england -> capital -> London" If you search for the key "germany" it has no results, if you search "london" it finds no results. The fluidity of the data is definitely a hindrance, if you wanted to use the dataset you'd have to already know what you are looking for to find the value.
- jitl 10y agoWhat sort of backend storage does this use?
- Arcsech 10y agoGiven that it's effectively key->key->value, I'm guessing Cassandra for the main backend. The data model fits very well, and it would give you the kind of scalability you would need for this sort of thing.
- smarx007 10y agoGiven how ridiculously basic the website is, I think you're right. It should have been a real triplestore with a SPARQL endpoint though.
- 0xmohit 10y agoYes, it would be interesting to get some insight into the tech stack.
- kristopolous 10y agoSo searches like "Linux" "California" "lincoln" "Hitler" "Disney" and "red" return 0 results
- outpan 10y agoUp until last week, outpan was strictly used for gathering data on product barcodes. We are in the process of adding data in other categories.
- outpan 10y agoI was driving so couldn't stay on the phone. The number of keys for different concepts is just very large (if not infinite for the sake of avoiding philosophical debates). It will take a some time to have enough data to cover even popular keys. Fortunately I have a lot of time :D
- GuiA 10y agoSignup page is blanking out for me. Really curious to try it. It's a really neat idea. Has this been done before? I've never seen anything like it. The full value of this would likely come from interesting, productive, insightful visualizations of the underlying graph that is being built. Questions that come to my mind: - What if you write bots that scrape Wikipedia, Twitter, etc. and output entries from semantic analysis performed on these sources? - If many people write such bots, how similar would the graphs be? What are the parameters that determine graph overlap? - Can you use this to tie in to the real world? "A key can be any unique string such as product barcodes, book ISBNs, email addresses, URLs, domain names, names, phone numbers ..." Interesting stuff there... A way to make this into an interesting social network is if people curated their own graphs, e.g. of books and webpages and favorite restaurants, and other people could browse these graphs in a read only mode. (perhaps they can clone it to their graphs or so and then start to edit them) The temporal aspect might be interesting too. Would there be value in seeing how my graph has changed over time? When I was in academia, none of the tools to keep track of the papers I read suited me. I could see this working well for this use case scenario. (if you could write a Prolog against this, would it have interesting properties?) I see that a submitter just posted 7630028603780 -> volume -> 75ml (the first number being probably the bar code of some beverage), that's neat. Someone else just posted the following entry: 5010438013621 -> ingredients -> Water, sugar, mixed fruit juices from concentrate 10% (Grape, blackcurrant, raspberry), acid (Citric acid), Vimto flavouring (Includes natural extracts of fruits, herbs, barley malt and spices), colouring food (Concentrate of carrot, hibiscus), acidity regulator (Sodium citrate), preservatives (Sodium benzoate, Potassium sorbate), Vitamin C, Sweeteners (Sucralose, Acesulfame K). Now what the site should be doing is converting this to something like: 5010438013621 -> ingredients -> water 5010438013621 -> ingredients -> sugar 5010438013621 -> ingredients -> vitamin C ... (Outpan person/people, I'm in SF, if you want to chat more around coffee, there's an email is in my profile)
- outpan 10y agoI think graph overlap is what actually determines what data is "accurate". There are currently bots writing to the database by people who are not connected and I have yet to see the overlap (since the key space is too large for the number of bots right now). I'm excited to see how that plays out. I will send you an email :)
- erikb 10y agoIs 55 million keys a lot for tracking "everything"? My expectation was that "everything" would need more than 55 trillion keys.
- outpan 10y ago"a journey of a thousand miles begins with a single step"
- yawniek 10y agodoes silicon valley now suddenly discover the semantic web and triplestores :O if you want really all the data: http://lod-cloud.net/ http://lod-cloud.net/ but still, neat project.
- outpan 10y agoSemantic web is too good of an idea not too iterate on constantly :)
- syats 10y agoThere's some more complex and curated projects out there, among them: https://www.wikidata.org/wiki/Wikidata:Main_Page https://www.wikidata.org/wiki/Wikidata:Main_Page
- singularity2001 10y ago+1001 for Wikidata compare https://www.outpan.com/data?k=germany https://www.outpan.com/data?k=germany to http://netbase.pannous.com/html/verbose/germany http://netbase.pannous.com/html/verbose/germany (wikidata mirror) or https://www.wikidata.org/w/index.php?search=Germany https://www.wikidata.org/w/index.php?search=Germany
- hjalle 10y agoNetbase/pannous seems to be cool but I guess it still got some work to do as it's suggesting Bonn as capital of Germany, unless I'm misinterpreting the statements. Kind of weird result from Google when you search for: "bonn capital germany". For me, it displays the wiki summary of Berlin but the link goes to Bonn. Does it usually work like that? Edit: It was the Bonn Summary of wikipedia
- singularity2001 10y agoBonn was the capital of Germany for many years.
- ivoras 10y agoShame it fails for the trap of natural language ambiguity. So on the front page I see "England -> capital -> London" and at a glance I thought the capital (as in money) flows from England and is accumulated in London.
- OJFord 10y agoI agree in general (and it's the fault of freely user-definable keys/attribute) - but that's not a great example, since the arrow aren't read as a direction of flow. Perhaps that reveals that a forward slash might be a better separator though - like a URI.
- firewalkwithme 10y agoThe header looks like fastmail :)
- 0xmohit 10y agoIs it possible to lookup all the attributes for a given value, say Trump?
- jrochkind1 10y agoSo it's RDF without the globally unique identifiers?