5 ms·
>Ever tried to look up some news from 12 years ago? I have a better one for you. Ever wondered why it's so hard? Why web protocols have nothing related to arch
by BrainVirus 4y ago
>Ever tried to look up some news from 12 years ago?
I have a better one for you. Ever wondered why it's so hard? Why web protocols have nothing related to archiving? Why web browsers are a hellscape for aggregating information over time in a meaningful way? Why this continues to be true, despite countless Microsoft and Google engineers writing all these heartfelt posts about knowledge?
If your answer is "because it's hard to implement" than you understand nothing.
- doliveira 4y agoFrom my admittedly limited understanding, the failure of the semantic web is one of mankind's biggest missed opportunities. Now the knowledge graph is just locked behind Google's neural network layers and only being used for ads.
- stjohnswarts 4y agoThe idea behind semantic web was inspiring and great, however it required considerable work on the part of people creating stuff for the web and that was never going to happen. Maybe it could have happened in some things like academia based or knowledge based websites, but on the larger scale it was doomed.
- hobofan 4y ago(Warning: Personal plug incoming) I fully agree, especially when it comes to the "semantic" part of the semantic web. Reusing and publishing ontologies that define those semantics always seemed like an afterthought of the semantic web, when it should be part of the foundation that things on the semantic web are built on. In most other parts that make up a website (JS and HTML) we figured out how to make reuse (mostly) work by replacing flimsy web references with package management. Ontologies never had something like that, and thus were stuck in an early 00s era of software/ontology development. Where I work, we are building Plow, a package manager for ontologies (https://github.com/field33/plow https://github.com/field33/plow) as part of our tech stack to improve that situation and allow people to build applications with large-scale stable semantics at the core. As part of building Plow we are aiming to make the process of creating and sharing ontologies easier and with that also lowering the barrier of entry to that domain.
- gnramires 4y agoMaybe something like a 'WikiInfo' (or another better name :) ), that contains a hierarchy of (potentially all) known pages and topics? I think the only way to tackle this problem is collaboratively and distributedly. You could add for example a 'Newspapers' topic, and then say 'The Springfield Times' and then have 'Articles by date', 'Articles by topic', etc. like a huge database (browsed hierarchically like "WHERE dates BETWEEN '20121211' and '20121213'", etc.). The primary datastructure could be a database, and users can add hierarchies as queries to the underlying database -- a collaborative index (in the literal sense, like a Homepage of the internet) is shown. Any unique 'object' (like a specific newspaper) gets an UUID and a row in the db. I don't know how modern dbs handle sparse data, but that'd definitely be a requirement (i.e. each object can have a handful of millions of possible properties, like publication date, location, author, colour, etc.).
- gnramires 4y agoI've looked online and there's WikiData[1], which I didn't know and looks very nice. Although it seems to be more of a plain database, not concerned with Indexing. It also doesn't seem to contain objects such as all newspaper articles (without the text body of course), I wonder if that would be accepted data. Maybe we could build upon WikiData as a backend and present a hierarchical index. As a humble suggestion, I'd divide all information in: (1) News (all news articles), (2) Publications (books, blogs, magazines, etc.) (3) Ideas and things (countries, planets, people, theories). Any object can belong to multiple categories/classes. I think it's not important that categories be perfectly devised, only that they contain all objects, and objects can be found reasonably well within them. The main point of objects would of course be a link to an actual web page that contains what you're looking for. Please, someone do it! (I have way too many projects right now) [1] https://www.wikidata.org/wiki/Wikidata:Main_Page https://www.wikidata.org/wiki/Wikidata:Main_Page
- xerox13ster 4y agoEdgeDB is half of what you describe on the database front. You can add nested queries with filters and calculated query types. Every item in the database is given a unique UUID and there's support for complex custom types and constraints, included calculated constraints I believe. You could have an Article type with a link to a Person as the author, and many more types of Works besides. You could find any Work by a Person traversing backlinks to find any object linked to that person. Any work where they contributed. Any social media link they posted. Then query that by a duration of time starting from a specific date. A specific place, if it has it, a group of sites and so on. If my understanding of it from my time playing with it is correct. I haven't experimented with too broad a dataset yet.