21 ms·
Show HN: I'm building a non-profit search engine
- yuhong 5y ago
- _xnmw 5y agoThis is a business model I've been thinking about: what if users earned credits for running a crawler on their machine? In other words, as much as I hate crypto scams, a "tokenized" search engine where the "mining" power was put to good use, i.e crawling and indexing.
- thebeastie 5y agoHow would you judge that they had actually done the work? The output needs to be verifiable.
- _xnmw 5y agoThere would have to be some aspect of centralized moderation, I suppose. This is beyond my knowledge: Is there a way to accept output only from signed binaries, so that we assume if X cycles of work were performed by a signed binary, then it is legitimate output?
- born-jre 5y agono not really may be some theoretical way using ZK-snarks or encrypted enclaves (intel sgx) but not practical. also probably does not work cz oracle/enclave also needs input(raw crawl data/ network http bytes) which has to be trusted. one way could be project could make unbreakable mining/crawling chip/box/os/anticheat_layer_with_vm and supply to everyone which has different levels of breakable-lity and complexity to build. One way is we send crawl_task to different to N random nodes and accept one that most similar? another way could be build messy network to solve messy problem. What we do is build reputation bashed graph network and you accept index from nodes you trust. so people will start un following misbehaving nodes. there is not universalroot view of network instead its dynamic and different from prospective of each node. or it could have one root view if we store reputation data in bchain and with some type quadratic voting to modify the chain. ? yeah Bitcoin showed us way to build mathematically secure system without any trusted party but it could do that cz problem it was solving is mathematically provable. Problem like collecting indexing crawl data you have to trust somebody.
- b3kart 5y agoWell they are in a sense: you can just do the task yourself. It’s expensive of course, so you can use methods applied to human labelling for ML: _periodically_ injecting tasks with known results and checking how trustworthy the party is, vending the task to multiple parties and aggregating results, blocking parties that make many “mistakes”, etc.
- hericium 5y agoDepending on who would be providing URIs. If the miner, they would be able to deliver any crap so content deliveries would have to be judged in some way and awarded differently. If the pool provides addresses to crawl, the miner could be given crafted/dedicated URIs from time to time and lack of delivery of Proof Of Crawl could result in a penalty chosen in a way rendering "cheating" unprofitable. But then fresh URIs have to come from somewhere.
- deleted 5y ago[deleted]
- GistNoesis 5y agoYou build a system based on trust but verify. If the output is the result of a known deterministic program on a known input, anyone can verify it. So they just have to sign their work. If later someone find that they have lied and provided a false result, they lose reputation/stacked coins. The second associated problem is how would one prevent them from appropriating the work of others that would just re-sign it. One way would be to allow the worker to introduce a few voluntary errors but have a secret joker that allow him to pass the challenge of a failed verification. One alternative way is based on data malleability. The worker pick a secret one way function and compute is F( data + secretFunction(data,epsilon) ) ~ F(data) and publish the values of the secretFunction(data,epsilon) but not the secretFunction. Only someone with knowledge of the secretFunction can make a claim on the work done. If there is a challenge only the real worker will be able to publish the secret of the secretFunction (Or use some zero knowledge proof to convince you they know it).
- charcircuit 5y ago>If later someone find that they have lied and provided a false result, they lose reputation/stacked coins. A web page isn't an immutable piece of text. It can change on every visit and it can sometimes returns errors.
- GistNoesis 5y agoThat's why you don't index the webpage but a snapshot of it. For example you index the commoncrawl archives, or some content addressable storage like ipfs or torrent file.
- charcircuit 5y agoWhat's the point of crawling the common crawl archive? It's pointless. You can simply download it.
- GistNoesis 5y agoI think crawling and indexing should be treated differently. Indexing is about extracting value from the data, while crawling is about gathering data. Once a reference snapshot has been crawled, the indexing task is more easily verifiable. The crawling task is harder to verify, because external website could lie to the crawler. So the sensible thing to do is have multiple people crawl the same site and compare their results. Every crawler will publish its snapshots (which may contain some errors or not), and then that's the job of the indexer to combine multiple snapshots of various crawler and filter the errors out and do the de-duplication. The crawling task is less necessary now than it was a few years ago, because there is already plenty of available data. Also most of the valuable data is locked in walled garden, and companies like Cloudflare make the crawling difficult for the rest of the fat tail. So it's better to only have data submitted to you, and outsource the crawling.
- zomglings 5y agoOne idea I have been kicking around is the idea of federating the indices, not the crawling. If every contributor maintained their own index, then you could reward contributors based on how many hits their index generated. This would open up the possibility of people maintaining indices for specialized topics that they were experts in, and give the federated search engine a shot at taking on Google.
- bspammer 5y agoYou don’t want to reward people for quantity, but quality. The cost of creating a new webpage is effectively zero, so if you attach an incentive for creating them you are doomed.
- zomglings 5y agoYou and I are talking about completely different types of people. For most people, the cost of creating and maintaining a website is high. This is why products like Wix and Squarespace exist (and are not cheap). I am thinking a simple dashboard where anyone could go and curate a list of content they find useful. They could share this with the world. The interface should be so simple that my parents could use it - and they aren't going to be putting up websites anytime soon.
- DarylZero 5y agoThis is what Yahoo was doing before Google took over. (But of course it wasn't a volunteer public benefit effort like you describe.)
- g105b 5y agoI'm very intrigued by this concept.
- Piezoid 5y agoYaCy is decentralized, but without the credit system. Some tokens, like QBUX, have tried to develop decentralized hosting infrastructure. I also have been wondering how this would play out with some kind of decentralized indexes. The nodes could automatically cluster with other nodes of users sharing the same interests, using some notion of distances between query distributions. The caching and crawling tasks could then be distributed between neighbors.
- _xnmw 5y agoYaCy is too slow for mainstream use. I believe the indices still need to be centralized, only index-building and crawling can be distributed.
- marginalia_nu 5y agoA big part of the problem I see with decentralized search is that you basically need to traverse the index in orthogonal axes to assemble search results. First you need to search word-wise in order to get result candidates, then sort them rank-wise to get relevant results (this also hinges upon an agreed-upon ranking of domains). That's a damn hard nut to crack for a distributed system. Crawling is also not as resource consuming as you might think. Sure you can distribute it, but there isn't a huge benefit to this.
- thebeastie 5y agoActually I have an idea for you: i think you can use cryptography to prove that an SSL session really happened. So you could prove indexing of HTTPS sites.
- thebeastie 5y agoI think the way this works is having the code to execute an ssl session encoded in a zkSnark. One of the zkSnark based blockchains is doing it.
- detaro 5y agoYou can prove that a TLS session happened, but nothing about its contents, so you can't really prove indexing.
- emrah 5y ago> To be honest, at the moment the algorithm doesn't return any sensible results for anything Is there an objective way to measure this? Do we just compare the output to Google or DDG? This seems like one of the many big hurdles in creating a competitive product in this space
- tibbar 5y agoIt’s really fast - nice job! Can you elaborate on the ranking algorithm you are using? It seems that this will become more important as you index more pages.
- daoudc 5y agoThanks! A really simple one for now: number of matching terms, and then prioritising matches earlier in the result string. But this is something I'm looking forward to working on properly when I get a bigger index.
- daoudc 5y agoI also want to incorporate a community aspect to ranking, allowing upvoting and downvoting of results. I've not yet figured out how to reconcile this idea with not having any tracking though. Perhaps a separate interface for logged-in users.
- foxfluff 5y agoOne ambitious project I've thought about over and over again over the years is search (and social sites / forums) where the votes, tags, and flags make a public dataset and users can manipulate their own weights (or even the ranking algorithm) to construct a "web of trust" that yields favorable results. This way you can escape spammers, powertripping moderators, and the tyranny of the hive mind; it doesn't matter if there's a large population of spammers, shills, and idiots upvoting crap because you set their weights to zero (or negative). In fact, that becomes a feature, because by upvoting crap, they generate a crap filter for you. If the weights are also public, then you can automatically & algorithmically seed your web of trust (simplest algo for sake of example: give positive weight to identities who upvoted and downvoted the same way you did) but you could still override the algo with manually set values if it gave too much weight to bad actors. Obviously this has privacy implications (all your votes and your network becomes public), and can generate a large dataset (performance challenge, how do you distribute it / give access to it?), so it's far from a trivial project. For the privacy angle, I'd start by keeping identities pseudonymous (e.g. a public key or random id -- you don't know who's behind the identity unless they blurt it out). Furthermore, I think it'd be useful to automagically split your actions across multiple identities so it's harder to link all your activity. I think the system should also explicitly allow switching identities, for privacy but also because sometimes you just want a different "filter bubble" which helps tailor the content you get to what you're looking for. Maybe the network that yields best shopping results isn't the same network that yields best cooking recipes or technical docs. With this model, everyone is a moderator and everyone can defer moderation to identities they trust, but neither the hive mind nor individuals have the ultimate power to dictate what you see. If you want to read spam or conspiracy theories, you just switch to your identity which upvotes such content and has positive weights towards other identities with similar votes. I doubt you're going to build this; I doubt people want this. I certainly want it. Maybe one day I'll try, but it probably won't work well without network effects (=reasonably large quantity of users). I just wanted to let you know about the idea because your project is inspiring and inspiring things inspire me to share ideas.. :)
- montebicyclelo 5y agoHow much compute, storage, and network speed would a minimal up-to-date web search engine need?
- iopq 5y agoIf you have to ask, you can't afford it
- dotancohen 5y agoFair point. How is it to be funded?
- daoudc 5y agoThe plan is to fund it through donations
- smt88 5y agoCan you make it a contributory database? I wouldn't mind "donating" my browsing history and page downloads to build the index and train the algorithm. You'd have to find a way to verify reputation to make sure no bad actors could contribute.
- aspenmayer 5y agoHow does Internet Archive verify dumps submitted by Archive Team and other groups? This may already be a solved problem. Not knowing their implementation details, I’m guessing it could be doable without reinventing much. An oracle could dispatch a P2P archive job to a pool of clients randomly assigned tasks, with both the first to archive and the first to validate being recognized by the swarm somehow, with periodic re-archiving and re-verification, rate adjusted by popularity of site and of search keywords.
- daoudc 5y agoYes, I'm planning to do something like this.
- amenod 5y agoOff-topic [0]: I would be very interested in an economic model that would work for such a search engine. Donations are fine, but (imho) it will take much more than that to keep the lights on, let alone expand... The "fairest" solution for both sides I can think of is ads which no not send tracking information, and are shown primarily based on search terms and country, or even other parameters that the visitor has set explicitly. Any other ideas on how to finance such an engine so that incentives are aligned? [0]: EDIT: off-topic because the page clearly states that this project will be financed with donations only.
- daoudc 5y agoWikimedia has an estimated $157m in donations this year. If we could get a small fraction of this amount we should be able to build something pretty good.
- klohto 5y agoGet real lol. Why would a general public care about you? Happy to donate but it won’t keep the light on. You’re serving a niche community
- daoudc 5y agoNiche for now, but I think a lot of people can see the value of search without ads.
- beachy 5y agoA journey of a thousand miles begins with a single step.
- oefrha 5y agoI wish you luck, but I mean, I use Google and I haven’t seen a search ad for what, a decade (okay, less than a decade considering iOS)? Most people who don’t want to see search ads can pretty easily find an ad blocker.
- 5y ago
- alexdowad 5y agoSome idle words from a passer-by: It would have been good if this project had a pronounceable name. "To Google" has entered the English lexicon as a verb, but I don't think anybody will ever say they "mwmbled" something.
- daoudc 5y agoIt's pronounced "mumble". I live in Mumbles, which is spelt Mwmbwls in Welsh.
- yetanother-1 5y agoNice, but still not very intuitive nor common for the grand public.
- danpalmer 5y agoSpelling it “mumble” wouldn’t be accurately pronounceable for most of the world, and billions couldn’t even read the letters. I get your point but I think we should normalise things that don’t come completely naturally for English speakers.
- jesprenj 5y agoMumble is already a group talk protocol. http://mumble.info http://mumble.info
- scottmcdot 5y agoIf it took off we all might start swapping e with w as a nod to our preferred search engine.
- wodenokoto 5y agoIn the early web 2, it was very in for things to be spelled unpronounceably. For the life of me I can only remember Twittr, but I wanna say Spotify also had an unreadable name in the early days.
- KarlKemp 5y agoThe central problem with this and similar endeavors: nobody is willing to pay what they are worth in ads. Let's say the average Google user in the US earns them $30/year. Are you willing to pay $30/year for an ad-free Google experience? Great! We now know that you are worth at least $60/year. That little thought experiment is true for many online services, from social networking to (marginally) publishing. But nowhere is it more true than for search results, which differ in two fundamental ways: being text-only, they don't bother me to anywhere near the degree of other ads. And, second, they are an order of magnitude more valuable than drive-by display ads, because people have indicated a need and a willingness to visit a website that isn't among their bookmarks. These two, combined, make this the worst possible case for replacing an ad-based business with a donation model. The idea mentioned in this readme that "Google intentionally degrades search results to make you also view the second page" is also wrong, bordering on self-delusion. The typical answer to conspiracy theories works here: there are tens of thousands of people at Google. Such self-sabotage would be obvious to many people on the inside, far too many to keep something like this secret.
- daoudc 5y agoTBF I don't think Google intentionally degrades results, but they have less incentive to improve the results.
- dazc 5y agoThe same way people don't intentionally break the law, they just overlook certain aspects of it when it suits them.
- dazc 5y agoDuck Duck Go is profitable despite not blanketing the first page with ads (just like google were once upon a time); you can have no ads at all if you like also. Do they make money in other ways, sure but not in a way that degrades the user experience. Are DDG results inferior, for 95% of users no.
- nyuszika7h 5y ago
- ortuman84 5y ago> Marginalia Search is fantastic, but it is more of a personal project than an open source community. And where's the community behind mwmbl project?
- daoudc 5y agoFeel free to email me if you want to be involved!
- Closi 5y agoHey, great project - the more competition in this space the better. To be honest, at the moment the algorithm doesn't return any sensible results for anything (at least that I can find), but I hope that you can find a way past this as it's a great place to have a project. I've included some search terms below that I've tried - I've not cherrypicked these and believe they are indicative of current performance. Some of these might be the size of the index - however I suspect it's actually how the search is being parsed/ranked (in particular I think the top two examples show that). > Search "best car brands" Expected: Car Reviews Returns a page showing the best mobile phone brands. then... > Then searching "Best Mobile Phone" Expected: The article from the search above. Returns a gizmodo page showing the best apps to buy... "App Deals: Discounted iOS iPhone, iPad, Android, Windows Phone Apps" > Searching "What is a test?" Expected result: Some page describing what a test is, maybe wikipedia? Returns "Test could confirm if Brad Pitt does suffer from face blindness" > Searching "Duck Duck Go" Expected result: DDG.com Returns "There be dragons? Why net neutrality groups won't go to Congress" > Searching "Google" Expected result: Google.com Returns: An article from the independent, "Google has just created the world’s bluest jeans"
- daoudc 5y agoThanks for the feedback! I'll take a look at your examples and see if I can improve the rankings.
- devoutsalsa 5y agoThe first thing I do to test a search engine is to search for my own username on various public sites to see if it can find me. It didn’t find me. But keep it up and I’m sure I’ll be in there eventually (or maybe a overestimate how interesting I am, hehe).
- rmbyrro 5y agoI get this is your usual testikg case for search engines, but if you've read their README you'd have seen its inaproppriate for the project at the current stage.
- legofr 5y ago> All other search engines that I've come across are for-profit. Please let me know if I've missed one! https://www.ecosia.org/ https://www.ecosia.org/ https://ask.moe/ https://ask.moe/ https://ekoru.org/ https://ekoru.org/ I remember seeing one more non-profit search engine on HN but can't seem to find it right now.
- daoudc 5y agoThanks, but these are not technically non-profit: "Ecosia is a search engine based in Berlin, Germany. It donates 80% of its profits to nonprofit organizations that focus on reforestation" [1] "80% of profits will be distributed among charities and non-profit organizations. The remaining 20% will be put aside for a rainy day." [2] "Ekoru.org is a search engine dedicated to saving the planet. The company donates 60% of revenue generated from clicks on sponsored search results to partner organizations who work on climate change issues" [3] [1] https://en.wikipedia.org/wiki/Ecosia https://en.wikipedia.org/wiki/Ecosia [2] https://ask.moe/ https://ask.moe/ [3] https://www.forbes.com/sites/meimeifox/2020/01/19/how-the-search-engine-ekoru-is-cleaning-up-our-oceans/ https://www.forbes.com/sites/meimeifox/2020/01/19/how-the-se...
- toper-centage 5y agoWhile Ecosia is not technically a non-profit, no one can sell Ecosia shares at a profit. > Ecosia says that it was built on the premise that profits wouldn’t be taken out of the company. In 2018 this commitment was made legally binding when the company sold a 1% share to The Purpose Foundation, entering into a ‘steward-ownership’ relationship. > The Purpose Foundation's steward-ownership of Ecosia legally binds Ecosia in the following ways: - Shares can't be sold at a profit or owned by people outside of the company and - No profits can be taken out of the company. https://www.ethicalconsumer.org/technology/how-ethical-search-engine-ecosia https://www.ethicalconsumer.org/technology/how-ethical-searc...
- gardenfelder 5y agoSeems like a reasonable approach is to use a b-corp where shareholders cannot sue for financial gains.
- discordance 5y agoIs there a test suite of expected search results that could be used with these sort of projects?
- ChemSpider 5y agoReally, I don't care if it is for-profit or not. Just a search engine with transparent ranking would be great. Ideally with explainable AI (XAI) that can tell me WHY is result A ranked higher than result B. I would even pay a monthly subscription to use it.
- challenger-derp 5y agoDo you yearn for explainability due to getting irrelevant search results? Is what you're searching for more specialized than what the public might consider general knowledge?
- wolfgarbe 5y agoA laudable effort. Two questions: 1. What is the rationale behind choosing Python as a implementation language? Performance and efficiency are paramount in keeping operational costs low and ensuring a good user experience even if the search engine will be used by many users. I guess Python is not the best choice for this, compared to C, Rust or Java. 2. What is the rationale behind implementing a search engine from scratch versus using existing Open Source search engine libraries like Apache Lucene, Apache Solr and Apache Nutch (crawler)?
- Faaak 5y agoPremature optimization is the root of all evil. Best to concentrate on the algorithm first, and then, maybe, improve it with a faster language. Apart from that, the misconception that "python is slow" should die :-)
- authed 5y ago"Apart from that, the misconception that "python is slow" should die :-) " Yeah it's not python that is slow, it's the interpreter.
- nulbyte 5y agoWhich interpreter? There are multiple. I found pypy to be quite reasonable; often faster than the standard C python interpreter.
- marginalia_nu 5y agoIn general, speed isn't the problem with search (at least the retrieval aspect), but memory efficiency is. Things like small object overhead and the ability to memory map large data ranges are extremely beneficial for a language if you want to implement a search index. But I agree, get it working first, then re-implement it in another language if it turns out to be necessary.
- wolfgarbe 5y agoThis is no contradiction, you can concentrate on the algorithm also in a faster language :-) https://benchmarksgame-team.pages.debian.net/benchmarksgame/fastest/python3-gcc.html https://benchmarksgame-team.pages.debian.net/benchmarksgame/... https://benchmarksgame-team.pages.debian.net/benchmarksgame/fastest/python3-java.html https://benchmarksgame-team.pages.debian.net/benchmarksgame/... https://benchmarksgame-team.pages.debian.net/benchmarksgame/fastest/python3-go.html https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
- deleted 5y ago[deleted]
- slmjkdbtl 5y agoAs a normal human I naturally typed in "fuck" in a new search engine and it led me to this article https://mathbabe.org/2015/06/22/fuck-trigonometry/ https://mathbabe.org/2015/06/22/fuck-trigonometry/ which I quite enjoyed!
- born-jre 5y agothere seems to be a lot of comment about some form of distributed trust/reputation bashed system.
- ZeroGravitas 5y agoI have this feeling that most of time I "search" for something I already know what I'm looking for, but google via firefox's omnibox, is just the fastest way to get there, even though it's a bit indirect. Are they getting paid for that, or am I costing them money in the short term, but they get to build up a profile on me to provide more effective ads later? I wonder if it's possible to take advanage of that type of search by putting a facade in front of the "search engine" and based on the search term and the private local user history, then go direct to a known site, or if it seems a search is needed, go to a specific search engine. This may open up opportunities for say program language specific search engines, or error messages from a program specific search, or shopping for X sites.
- yellowsir 5y agoif u set duckduckgo as your default search provider, you can use bang in the omibox. also you can toogle between local-area or global search. https://duckduckgo.com/bang https://duckduckgo.com/bang e.g. !yt !osm !gi
- medstrom 5y agoI bookmark every site I might possibly want to revisit - make a habit of Ctrl+D. They're totally unsorted, but the key is to wipe the regular history on exit, leaving only the bookmarks as source material for completion. That way I can type something in the url bar and get completion to interesting sites. The url bar (or omnibox) matches on page title as well as the actual address, so it's easy, and always faster than a search engine.
- bluecatswim 5y agoMost wikis or resource/documentation sites have a local search bar on their homepage, Firefox has a feature where it lets you add a search keyword for that specific site. So if you add, say, pydocs as a keyword for docs.python.org you can do "@pydocs <query>" it looks up the query on that page.
- amelius 5y agoWhat literature did you use to obtain suitable algorithms for search/NLP?
- tomxor 5y ago> We plan to start work on a distributed crawler, probably implemented as a browser extension that can be installed by volunteers. Is there a concern that volunteers could manipulate results through their crawler? You already mentioned distributed search engines have their own set of issues. I'm wondering if a simple centralised non-profit fund a la wikipedia could work better to fund crawling without these concerns. One anecdote: Personally I would not install a crawler extensions, not because I don't want to help, but because my internet connection is pitifully slow. I'd rather donate a small sum that would go way further in a datacenter... although I realise the broader community might be the other way around. [edit] Unless, the crawler was clever enough to merely feed off the sites i'm already visiting and use minimal upload bandwidth. The only concern then would be privacy. oh the irony, but trust goes a long way.
- daoudc 5y agoYes, that is a concern. I'd probably worry about it if and when it started happening, however.
- Schiendelman 5y agoIf you wait until then, it may become too late to mitigate. Unless you have a plan to remain in complete control of who contributes.
- bryguy32403 5y agoFamous last words
- altdataseller 5y agoThere are already loads of browser extensions that do a lot of screen scraping of all the sites you visit, without you even realizing it,
- tomxor 5y agoI can imagine, but i only use one, ublock
- btdmaster 5y agoAny reason to prefer GPLv3 over AGPLv3? It might be useful to use the latter so that distributing it over a network requires distributing modifications as well.
- daoudc 5y agoGood suggestion, thanks.
- deleted 5y ago[deleted]
- juliushuijnk 5y agohere's my open data attempt from a couple of years ago: http://www.charitius.org http://www.charitius.org goal was/is to include all charities in the world, based on open data, and open software. It's been on pause for a while, but still works, and open for new sources of data to incorporate.
- prohobo 5y agoThank you for this, but I feel like you should make a profit and this is currently a missed opportunity to use web3 principles to do that. Free and open software is a great ideal, but the reality is that people need money to live - and ads are the way to make that money on web2 platforms, which is why Google is in such a sad state. Why not do something similar to Brave? You can add tokenomics to the search engine and make money while keeping it 100% useable and open source.
- everydaybro 5y agoexactly, free and opensource is amazing, but the developer should also earn money to live, maybe add donations, or maybe add some web3 features
- deleted 5y ago[deleted]
- astoor 5y agoHow would web3 and "tokenomics" solve search? If the underlying problem is that spamdexers are given a profit incentive to game search engine results with low effort and low quality content, does it make a difference whether the profit incentive is via pages generating advertising revenue or pages generating cryptocurrency revenue?
- prohobo 5y agoI don't really want to address your questions; this has been talked to death. How do you think tokens could help? Maybe they'd create incentives for users to use the engine. Maybe they'd open a market for said tokens so the dev can extract currency out of it. Maybe tokens can be a meta-system on top of the search engine, so that the search functionality can be left to solving the search problem without interference. Do you think there's an alternative to tokens in order to fund the project without degrading the search algorithm? If so, I'm all ears. In fact, I think we all are. Please enlighten us. But you haven't proposed a solution while crypto devs have been working on one since 2008.
- marcodiego 5y agoNon-profit search engines are needed. It will probably still be vulnerable to SEO but will more likely be resistant to become corrupt by the interest of "investors".
- quantum2021 5y agoTwo big things that annoy me about google: 1. They somewhat get around this with their maps feature, but their regular search doesn't actually search by area; you always get national websites that optimize the best. That would be a nice feature to have starting out without having to type in the specific area you're looking for. 2. Search results for hotels that actually work! Not only if they're set up on OTA's! This could actually get your search engine some traction as the search engine to go to when making travel plans which would give you a nice niche to start out in.
- deleted 5y ago[deleted]
- tmnstr85 5y agoGuideStar is the veteran in this space. I agree that doing this based on web scrapes and robots.txt is probably going to be pretty tough to get quality results. GuideStar always sells there product on the premise that they're vetting the financial statements of non-profits of best results. The real money might be on figuring out a way to scale reading and classifying non-profit financials - then see if you can quality control using a set of patterns.
- AlphaWeaver 5y agoThis isn't a search engine for nonprofits, it's a search engine that's designed to be run by a nonprofit one day.
- blueatlas 5y agoOr perhaps a for-profit company running this search engine and not garnering profit from it by directly changing search results, or having a different model for profit that does not affect search rankings.
- gkasev 5y agoCongrats on the mvp path you took to lunch your product. Generally, I think that there is a place for other variations of web search, be it in the way you crawl or perhaps how you monetize. I genuinely believe that it is really hard to build a general purpose search engine like DDG, Google and the like, but you can build a fairly good niche search engine. I'm particularly fond of the idea of community powered curation in search. Just today I lunched my own take on a community driven search engine - https://github.com/gkasev/chainguide https://github.com/gkasev/chainguide. If you like to bounce ideas back and forth with somebody, I'll be very interested to talk to you.
- marban 5y agoI've recently built one just for business news w/ obligatory zero-tracking. (https://yup.is https://yup.is)
- deleted 5y ago[deleted]
- freediver 5y agoCongrats! Very nice to see results being lightning fast, I am getting 100-120ms response with network overhead included and that is impressive. The payload size of only 10-20kb helps immensely, good job! I've built something similar called Teclis [1] and in my experience a new search engine should focus on a niche and try to be really, really good at it (I focused on non-commercial content for example). The reason is to be able to narrow down the scope of content to crawl/index/rank and hopefully with enough specialization to be able to offer better results than Google for that niche. This could open doors to additional monetization path, API access. Newscatcher [2] is an example of where this approach worked (they specialized on "news"). [1] http://teclis.com http://teclis.com [2] https://newscatcherapi.com/ https://newscatcherapi.com/
- gravypod 5y agoIf you filed to become a non-profit could people "donate" their engineering time as a tax write off? If you find out the legality of something like this and make it easy to do that could inspire a lot of collaboration on the project and I can see a bunch of other areas (outside of search) where services could be provided like this. I'm also sure having a non-profit would also make it easier to find cheap hosting which is a large part of the cost there.
- aantix 5y agoHow does someone economically store tens if thousands of terabytes of data needed for the indexes of a large scale search engine? And have the large server instances (lots of ram)?
- supernovae 5y agoI tried this back in 2006 - mozdex (only a wikipedia article survives) - it’s not cheap. I was a fan of lucene which led to nutch and eventually hadoop. so lots of servers running hdfs doing map reduce jobs to compile and update indexes. No one in the end cared about open search… duck seems to do alright under the guise of security but most non major searches are just meta searches these days because of economies of scale being highly disadvantageous to any upcoming search - and people just don’t search like they used to either. i was spending 2500 a month just on indexers and had query traffic taken off my costs would have shot through the roof since you want query nodes to all be in memory cache and that was expensive back then. today i would have used some modern in memory distributed doc dbs instead of query masters with heavy block buffer caches. i learned a lot but lost my shirt :)
- aantix 5y agoI think I would just start with indexing the latest Common Crawl WET files along with a server and a couple of Exos X20 20TB drives. Wondering out loud if there's a crawl that will produce real-time updates via common news outlets..
- supernovae 5y agoCrawling is the easy part. The 20tb drives would be too slow for query though. Query really needs everything in memory to be fast enough.
- marginalia_nu 5y agoWhy would you need that much data? The average website has maybe 10kB worth of textual information without compression. To get tens of thousands of terabytes of data, you'd need to index of the order 10^12 websites. That seems a bit much.
- champagnois 5y agoIf I were to work on building a search engine from scratch, I would probably approach this from these directions: (1) Investigate if running a DNS server will help me get a more robust picture of what websites exist. (2) Investigate if supplying a custom browser would help me to leverage client PCs to do the crawling / processing for me. (3) Investigate if there is any point in building a search engine with the data gathered in a non-proffit way... Non-proffits are not as sustainable as for proffit corporations.
- WheelsAtLarge 5y agoMake it open source and syndicate it. The goal is to get people to contribute both resources and code. Think about the Shopify as the model. Where many people contribute to create a huge shopping place. People care about their shop only but ultimately they create a useful shopping area. Also setup a foundation to guide its development and be able to hire a management team. The real challenge is not the code development but setting up an organization that will outlast all the challenges that will appear. Wikipedia is the model to follow.
- kova12 5y agoWhat do you do in order for your crawler to not accidentally weer into some naughty-naughty site and yield you a visit from your friendly FBI squad? That concern why I decided to stay away from yacy
- ChuckMcM 5y agoOkay, the cynical quip is "All search engines other than Google's are 'non-profit'." :-) But the reasons for that won't fit in the margin here. Building search engines are cool and fun! They have what seems like an endless source of hard problems that have to be solved before they are even close to useful! As a result people who start on this journey often end up crushed by the lack of successes between the start and the point where there is something useful. So if I may, allow me to suggest some alternatives which have all the fun of building a search engine and yet can get you to a useful place sooner. Consider a 'spam' search engine. Which is to say a crawler that you work to train on finding spammy useless web sites. Trust me when I say the current web is a "target rich environment" here. The purpose would be to not so much provide a search engine in total here, as it would be to provide something like the realtime black hole list did for email spam, come up with a list of URLs that could be easily checked with a modified DNS type server (using DNS protocol but expressly for the purpose of doing the query 'Is this URI hosting spam?' in a rapid fashion. There are two "go to market" strategies for such a site. One is a web browser plugin that would either pop up an interstitial page that said, "Don't go here, it is just spam" when someone clicked on a link. Or a monkey-script kind of thing which would add an indication to a displayed page that a link was spammy (like set the anchor display tag to blinking red or something). The second is to sell access to this service to web proxies, web filters, and Bing which could in the course of their operation simply ignore sites that appeared on your list as if they didn't exist. You will know you are successful when you are approached by shady people trying to buy you out. Another might be a "fact finding" search engine. This would be something like Wolfram Alpha but for "facts." There are lots of good AI problems here, one which develops a knowledge tree based on crawled and parsed data, and one which answers factual queries like 'capital of alaska' or 'recipe for baked alaska'. The nice things about facts is they are well protected against the claim of copyright infringement and so people really can't come after you for reproducing the fact that the speed of light is 300Mkps, even if they can prove you crawled their web site to get that fact.
- daoudc 5y agoUpdate: there's been interest from a few people so I've started a Matrix chat here for anyone that wants to help out or provide feedback: https://matrix.to/#/#mwmbl:matrix.org https://matrix.to/#/#mwmbl:matrix.org
- deleted 5y ago[deleted]
- rascul 5y agoIt looks interesting. However, the results appearing so fast as I type, and changing just as fast as I type more, makes it seem like it's flickering and it's painful on my eyes. Perhaps a slight delay and/or a fading effect as the results appear would be a bit easier for me to look at.
- daoudc 5y agoThanks for the feedback! Yup, a delay is on the to-do list.
- gryster 5y ago
- reacharavindh 5y agoThe idea of building a distributed crawler that runs on user’s browser sounds fascinating!! Now that is better than the user burning power to mine bitcoins solving stupid puzzles.
- lovefeature 5y agoNice,i'm building service engine https://beforedo.com https://beforedo.com , facebook alternative: https://alovez.com https://alovez.com instagram alterative : https://www.snapfeel.com/ https://www.snapfeel.com/