10 ms·
YaCy, a distributed Web Search Engine, based on a peer-to-peer network
- DrDroop 3y agoI once went to a workshop on a Sunday morning at the local makerspace to listen to someone talk about some kind of distributed search engine or something like that. One of the developers came from (I think) Germany to explain this to us the centralized sheeple. He just gave a demonstration of the thing, like here is the box you type stuff and here are the results. When I started to ask questions about how it worked an all he sort of acted annoyed saying it was all too difficult to explain. This was more than ten years ago, and yes I am still angry about it.
- ssijak 3y agoAt the core it was probably based on peer to peer distributed hash tables, so here you go, read the source https://pdos.csail.mit.edu/~petar/papers/maymounkov-kademlia-lncs.pdf https://pdos.csail.mit.edu/~petar/papers/maymounkov-kademlia...
- belter 3y ago160 bits ought to be enough for anybody :-)
- albert180 3y agoIt's probably him YaCy is made by a German Dude
- synctext 3y agoImpressive 20 year project by one key developer. See 20 year post in German by YaCy founder: https://community.searchlab.eu/t/yacy-vor-20-jahren/1543 https://community.searchlab.eu/t/yacy-vor-20-jahren/1543
- ssijak 3y agoLong time ago I worked for a startup called Wowd which built distributed search engine. It was acquihired by Facebook. On of the biggest issues was how to entice people to download and run the client/node. I half wondered afterwards if slapping some crypto on top of it which would be mined by running the node and providing resources would help. My gut says easy yes, but my mind grimace at the abomination.
- zoklet-enjoyer 3y agoWe have proof of stake now. The nodes could be run by the chain validators and they get a cut of the staking rewards. Look up how proof of stake works on the Cosmos chain. You could totally do this and I bet it would take off, at least in that section of the Internet that's into Cosmos/Tendermint chains. I'd use it
- ssijak 3y agoI was definitely thinking of some kind of proof of stake, not proof of work.
- zoklet-enjoyer 3y agoHahaha one downvote. I love to see it
- worksonmine 3y ago> but my mind grimace at the abomination Why would that be an abomination? It's a perfect use-case. Like you noticed people need incentives to volunteer their hardware. If you hate crypto because it's crypto you can just use fiat instead.
- komali2 3y ago> Like you noticed people need incentives to volunteer their hardware. I wonder if this is because "volunteer your hardware" projects sometimes involve someone making else money, and if someone else is making money but not you, why should you donate your hardware? For the truly libre "hardware donation" projects, they seem to be doing ok without financial incentivization. What immediately comes to mind is the petabytes of data flying around on peer to peer systems through torrenting. I know people that spend thousands of dollars a year on upkeep and upgrades for what are essentially super seedbox homelabs (I'm one of them too :P ) There's also communities like soulseek where people keep TBs of music up, often seeking out rare tracks to make available to the community for free. There's folding@home and seti@home, and I'm sure other similar projects I haven't heard of, where people donate cycles just for the common good. folding@home is a great example because we can directly compare the people that are "incentivized" to participate with bananocoin, a cryptocurrency rewarded based on work cycles in folding@home. You can see all bananocoin miners here under the banano.cc team: https://stats.foldingathome.org/ https://stats.foldingathome.org/ That team is in first place for work completed, however are only just surpassing the linus tech tips team, and not to mention compared to a bunch of other teams (and private "donors") they're a very small % of work completed for folding@home So therefore I disagree that people "need" incentives, there just needs to be no, erm, disincentives, if that's a word.
- rasulkireev 3y agoLove it. Super easy to self host and use. Now I have a personal Google!
- maxloh 3y agoSee also: Presearch, another decentralized search engine, claimed that it will be open source. No source code available at the moment though. https://presearch.com/ https://presearch.com/
- b2bsaas00 3y agoCould this be used for a Torrent search engine?
- fddrdplktrew 3y agoif it is not censored, probably?
- worksonmine 3y agoRecently there was a distributed tracker on the front page. Probably more what you're looking for.
- feverzsj 3y agobtdig is still alive.
- qingcharles 3y agobtdig has the data, but its search is subpar :(
- vGPU 3y agoHas it gotten any better recently? I run a node but I haven’t actually used it as a search engine in a while, as I found the result quality to be exceedingly poor.
- rahen 3y agoI remember trying it for a while in 2012, but the results were essentially worthless, probably because there were so few nodes/crawlers back then. I guess the more users there are, the better the results.
- viraptor 3y agoAlternatively, ignore the public network (it's still useless) and run it as your own crawler. Seed it with your browsing history, some aggregators like HN, your favourite RSS feeds, etc. and you'll be good.
- WarOnPrivacy 3y ago> I remember trying it for a while in 2012, but the results were essentially worthless, I had mine crawling gov, mil, etc sties for pages that Google was starting to delist back then. Inbound requests were heavy with porn until I tweaked - IDK, something.
- Brian_K_White 3y ago"until I tweaked - IDK, something." omg so much this. I got an instance going in a truenas core jail, freebsd and using freebsd java not a linux vm or linux abi compatibility. had to make my own rc script. Then had to mess with the disk & ram settings to get it to run for more than a day. But the settings are not actually explained at all and whatever they do, they definitely don't do what their names and worthless tooltips say they do. It seems to be running now indefinitely without killing either itself or the host, in full p2p mode, but I really have no idea why it's working, or really for sur if it actually is fully. I changed "idk, something" And I don't use it for search myself so far. Maybe some day but for now I'm paying for kagi. I just like the idea and want it to be a thing, and it seemed a little less "invite a world of shit and attention onto my ip" than running say a tor exit or something. Maybe only a bit less but I'll see how it goes and react if I need to.
- RGBCube 3y agocurl failed to verify the legitimacy of the server and therefore could not establish a secure connection to it. To learn more about this situation and how to fix it, please visit the web page mentioned above. Can't seem to access the page.
- gonesilent 3y agoInfrasearch / Gonesilent sold to Sun turned into project JXTA and died.
- mdaniel 3y agoWhile trying to read more about it, turns out there's an O'Reilly book, too: https://www.oreilly.com/library/view/jxta-in-a/059600236X/ch01.html https://www.oreilly.com/library/view/jxta-in-a/059600236X/ch... and there's also this https://wiki.wireshark.org/JXTA https://wiki.wireshark.org/JXTA (I'm guessing those specification links are in wayback but I didn't chase them)
- charcircuit 3y agoAre the results still being gamed by sites using content keyword stuffing? The last time I used it the searching and ranking technology felt like they were 40 years behind state of the art.
- liotier 3y agoIn distributed indexing, spam management seems a much bigger problem than the indexing itself.
- boyter 3y agoI actually half wrote a RFC of a spec and 2 implementations of a federated search last year. Rather than do the disturbed hash table that yacy does. I wanted results to be re-rankable by the peers by sharing the scores that went into them. The idea being with a common protocol based on the ideas of ActivityPub you could get peers of searches working together to hopefully surface interesting things. Something I should probably finish and publish at some point. It worked to the hundreds of peers I tested. The reason I mention this is because I wanted to also add a front into yacy which tuned out to be harder than I expected. It’s a wonderful project and you can find great stuff through it but the way the peers return results sometimes it’s hard to find it again. It’s also not quite as hackable as I would have hoped at the time probably due to he project age. I still think there is value in it though and I’d love to see yacy have its protocol explained as an apex so people could,build implementations in other languages more easily.
- detourdog 3y agoI remember the first days of gopher browsing were like that. Gopher browsing to me was like swinging on vine to vine. The trick was remembering/documenting where each vine went.
- arboles 3y agoSort of hijacking the thread to ask, can YaCy or similar, be an alternative to Google's Programmable Search Engine? All I use it for is limit a search to a medium-sized list of domains. The aspect that makes running a search engine difficult on your own is lack of resources for crawling, I expect. But since I only care about a small list of domains, could I ditch Google's and run my own crawler like YaCy?
- gtirloni 3y agoIs that the deceased code search tool? You could run Sourcegraph and import/sync those repositories. Or you could run your own ElasticSearch/Melisearch and crawl the websites yourself (if you're interested in things other than git repositories).
- arboles 3y ago> Is that the deceased code search tool? No, it's Programmable. Though it's not actually programmable. I should've written Custom Search Engine instead, that's also a name for it. cse.google.com - It's quaint that past the modern landing page, when using the search portal today, you still get some outdated iteration of Google UI design. It's used, for example, for making OSINT searches.[0] Or at some point by at least one Wikipedia editor for a custom list of Reliable Sources for Anime & Manga.[1] [0] https://www.osintme.com/index.php/2020/09/28/ https://www.osintme.com/index.php/2020/09/28/ [1] https://gwern.net/me#wikis https://gwern.net/me#wikis
- anthk 3y agoUgh, Java. I'll wait for something like i2pd does for I2P, something called yacyd either in c, c++ or golang.
- ravenstine 3y agoWhat's your objection to Java?
- anthk 3y agoHigh CPU and RAM usage.
- WarOnPrivacy 3y agoYacy's still around. Nice. After a year or two of hosting a Yacy instance (2014?) I started winding up on some general (probes, etc) blacklists. I also host a small mail server and I was getting mail returned. I'd force an IP swap and a few weeks later it'd be the same. I had to let Yacy go.
- 1oooqooq 3y agoSo that is how they block a people's search/crawler. Didn't thought they would use the most complicated method. They also use block lists to add every single TOR node (even if not an exit) and every VPN under the sun (except for streaming, because, why would them, that's why they exist)
- WarOnPrivacy 3y ago> So that is how they block a people's search/crawler. It seems less that was the intent and more they blacklist IPs with behavior they find annoying. They're super general lists. > They also use block lists to add every single TOR node This annoyed the crap out of me. The stupid Dan list guy made it as easy as he could to lump low risk bridges in with high risk exit nodes.
- renegat0x0 3y agoThere are already many project about search: - https://www.marginalia.nu/ https://www.marginalia.nu/ - https://searchmysite.net/ https://searchmysite.net/ - https://lucene.apache.org/ https://lucene.apache.org/ - elastic search - https://presearch.com/ https://presearch.com/ - https://stract.com/ https://stract.com/ - https://wiby.me/ https://wiby.me/ I think that all project are fun. I would like to see one succeeding at reaching mainstream level of attention. I have also been gathering links meta data for some time. Maybe I will use them to feed any eventual self hosted search engine, or language model, if I decide to experiment with that. - domains for seed https://github.com/rumca-js/Internet-Places-Database https://github.com/rumca-js/Internet-Places-Database - bookmarks seed https://github.com/rumca-js/RSS-Link-Database https://github.com/rumca-js/RSS-Link-Database - links for year https://github.com/rumca-js/RSS-Link-Database-2024 https://github.com/rumca-js/RSS-Link-Database-2024
- fsflover 3y agoBut which of those projects are distributed and FLOSS?
- legrande 3y agoAlso these: https://swisscows.com/en https://swisscows.com/en https://search.disconnect.me/ https://search.disconnect.me/ https://www.ecosia.org/ https://www.ecosia.org/ https://metager.org/ https://metager.org/ https://searx.space/ https://searx.space/
- ColinHayhurst 3y agohttps://www.mojeek.com/ https://www.mojeek.com/ self-disclosure, mojeek team member
- wongarsu 3y agoTo be fair, of those only Apache Lucene predates YaCy. YaCy is very mature, but in terms of relative popularity for general web search probably peaked around 15 years ago.
- buffalobuffalo 3y agoI ran YaCy for a while, but not as a node on their distributed search index. I just ran it as a search engine for all my own bookmarks. Unfortunately I never found a particularly good way of getting bookmarks into the system. So eventually I shut it down. Cool idea in theory though.
- justusthane 3y agoI have plan that I haven’t implemented yet, but I want to route all my outbound internet traffic through a Squid reverse proxy, which will in turn add every visited URL to YaCy (except for domains I choose to exempt). That way I’ll have a fully searchable index of every website I ever visit, which will hopefully solve the “Oh shit, what was that one website I found about X two months ago?” A potentially easier thing to do would be create a bookmarklet that adds the current page to YaCy.
- buffalobuffalo 3y agoYeah. Bookmark indexing was my original goal. But yacy doesn't have a great interface for that. Doable with some work, but not something i wanted to sink too much time into.
- mdaniel 3y agorelevant: https://github.com/ArchiveBox/ArchiveBox#readme https://github.com/ArchiveBox/ArchiveBox#readme and https://github.com/Rhizome-Conifer/conifer#readme https://github.com/Rhizome-Conifer/conifer#readme (nee "webrecorder/webrecorder")
- fortran77 3y agoRelated to this — I’d love to see individuals making web pages again, and federated search engines indexing them. People don’t make their own hobby or fan or art websites anymore, and I think that’s partly because nobody will every find them with the big search engines.
- emrah 3y agoI think it would be nice if the search results were "distributed" rather than deterministic. So when i enter the same keywords, let's say there are 50 pages each of which would be equivalently good result for the search, rather than one page "winning", the search engine would alternate the winner among the many possibilities
- jrussbowman 3y agoNice to see search projects are still popping up. After a move, family life taking over and me getting more interested in Unreal Engine, my poor search engine is now more of an experiment in seeing how well it runs while basically on life-support maintenance updates I do. Starting to think I honestly should just take it down and save my $50 a month I spend maintaining it. But I'll post it in a hacker news comment and maybe you all will give it enough traffic I can get excited about it again, lol https://www.unscatter.com https://www.unscatter.com
- jrussbowman 3y agoAnd for my immature moment of the day, the above comment was comment #69
- fho 3y agoI've been using several times over the last decades and never got good results. I think one instance is still running on my old computer at uni :-)
- dredmorbius 3y agoPreviously: YaCy – your own search engine | https://news.ycombinator.com/item?id=32597309 https://news.ycombinator.com/item?id=32597309 | 2 years ago | 93 comments YaCy: Decentralized Web Search | https://news.ycombinator.com/item?id=22246732 https://news.ycombinator.com/item?id=22246732 | 4 years ago | 41 comments YaCy – The Peer to Peer Search Engine | https://news.ycombinator.com/item?id=17089240 https://news.ycombinator.com/item?id=17089240 | 6 years ago | 3 comments YaCy: a free distributed search engine | https://news.ycombinator.com/item?id=12433010 https://news.ycombinator.com/item?id=12433010 | 8 years ago | 24 comments YaCy: Decentralized Web Search | https://news.ycombinator.com/item?id=8746883 https://news.ycombinator.com/item?id=8746883 | 9 years ago | 29 comments YaCy takes on Google with open source search engine | https://news.ycombinator.com/item?id=3288586 https://news.ycombinator.com/item?id=3288586 | 12 years ago | 17 comments
- treprinum 3y agoIs it worth dedicating 1-2 low power NUCs (4-8 core) to this on a 250MBit/s connection? Or does it need beefier CPUs/network?
- nairboon 3y agoIf you run YaCy with docker and it is still a junior peer, does the search return results from the global index or just the one that appears to be 'preinstalled'?
- Geranen07 3y ago[dead]