4 ms·
Your landing page is missing any kind of "evidence" that it is scaleable, low-latency or high throughput. Also if you are sharing on predicate you will end up
by jerven 11y ago
Your landing page is missing any kind of "evidence" that it is scaleable, low-latency or high throughput.
Also if you are sharing on predicate you will end up in big trouble. Predicates in most RDF datasets are not at all evenly distributed, tending more towards extreme value distributions. e.g. in UniProt the most common predicate has 2,419,000,171 occurrences, the least 1!
Also if you are going to benchmark can I suggest the rather good LDBC ones[1]. Even if for marketing reasons you don't want them public they are good to show where you can improve.
[1]http://www.ldbcouncil.org/ http://www.ldbcouncil.org/
- mrjn 11y agoThe way data is sharded across machines, is done via predicate. So, in this case, the 2.4 billion occurrences only leads to a data of 2.4G * 8 bytes (uint64 id) ~ 20GB of data for us, on one machine; which is pretty manageable. Furthermore, each predicate could be further sharded to fit on multiple machines; such would be the case with say the friend predicate in Facebook. It's hard to "prove" on a landing page, without going into design details, that you truly are those things. We do have a demo, with 21 million RDFs from real world data from Freebase, so you could play with the database, and get a feel for it. We'll look into LDBC. Thanks for the pointer.