5 ms·
I remember https://yacy.net/ https://yacy.net/ but the big problem of this project was java and had not implementations in others languages. I mean it as imagin
by mdtrooper 3y ago
I remember https://yacy.net/ https://yacy.net/ but the big problem of this project was java and had not implementations in others languages. I mean it as imagine torrent was only in perl.
- marginalia_nu 3y agoYaCy's big problem is that distributed search is a bad idea that will never perform well. Search is as fast as the data is local.
- kristopolous 3y agoThere was an effort in the early 90s to have search as a protocol so you could have a query and then select the domains you want to run it on and return an aggregate result. It was 100% abandoned and I think that's a mistake. It'd be nice to explore some of those ideas again
- marginalia_nu 3y agoI think a big part of the problem is that domains in isolation don't provide the best search results. Out-of-band information like (global) anchor texts or click data makes search perform so much better. If I want to learn how to do an INNER JOIN in MariaDB, this is the authoritative source: https://mariadb.com/kb/en/join-syntax/ https://mariadb.com/kb/en/join-syntax/ The problem being that INNER JOIN isn't particularly important to that page using most IR measures of importance, it's also primarily in a <code>-block which is typically further de-prioritized. To learn that this is an important link, you need to look outside of mariadb.com.
- kristopolous 3y agoThere's more to it than that. What if instead of crawling the php generation of database rows with a bunch of cruft, the administrator published some kind of schema with scraping and querying rules and you could alternatively make a single call to capture all of the data in a sematic schema. You can still do all the stuff you're talking about but it could make search more coherent. An entry for that humans and an entry for the computers. You can't trust everybody like this sure, but say imdb, discogs, wikipedia, all of which provide database dumps anyways (eg: https://datasets.imdbws.com/ https://datasets.imdbws.com/). That's what I'm advocating for revisiting. Lots of legit sites such as universities, newspapers, public records offices... You could even have a search toggle "screened sources" or whatever for the ones that make the cut
- teddyh 3y agoWasn’t this what the Semantic Web was supposed to enable?
- kristopolous 3y agoThat was more an ontological web. That's a different project which I totally support but this time through ML
- teddyh 3y agoThen, DBpedia might be more like what you’re after?
- teddyh 3y agoYou’re thinking of WAIS, I believe: <https://en.wikipedia.org/w/index.php?title=Wide_area_information_server&oldid=1170551140 https://en.wikipedia.org/w/index.php?title=Wide_area_informa...>
- kristopolous 3y agoYes!