26 ms·
The world needs a non-profit search engine
- mmazing 4y ago> The technology for organizing the world’s knowledge should be owned by everyone. This is what really nails it for me. There's far too much black box in pretty much every major search engine out there. Maybe it's by design "so that people can't game it". Even so, it's not working very well. I'm excited for the next 10 years to see what we (humans) come up with to solve the state of the internet, because something's gonna give at some point.
- sen 4y agoWhy though? 99.9% of people don’t care, and the internet (Facebook, instagram, TikTok, YouTube, gmail, the occasional Wikipedia) work fine for them. The overwhelming vast majority of people just don’t care about the internet outside those major few silos, so as far as “humanity” is concerned the internet is working as intended. It pisses me off what’s become of the internet, but I don’t personally see it changing.
- chaps 4y agoThis attitude is exactly why things won't change. .1% of the world is still 7 million people.
- chrischen 4y agoI guess it’s a good thing 99% of people don’t run the world. If products were dictated by the majority we wouldn’t have macOS, nor even iOS which is technically a minority by raw users. Obviously some users use more intensely than others, which is why catering to power users is a totally viable and legit strategy.
- lovemenot 4y agoYou are right about people's current opinions, but seem to be assuming that given a better option, people would continue to make sub-optomal choices. If so, I don't agree. Your small minority's job is to deliver those alternatives, and to feed the flames while the rest of the world makes the transition. Which they will do because the next thing is clearly so much better. Have more faith in most of humanity and doubt yourself and similar others for having failed so far to disrupt this industry with better technology.
- InCityDreams 4y ago>.....to see what we (humans) come up with to solve the state of the internet, because something's gonna give at some point. We (humans)? Are the fish up to something? Never did trust fish, especially regarding SEO stuff.
- 0xMatt 4y agoanother thing you could offer is really nice clean themes. I've paid for more than one app in my time JUST to get dark mode lol. obviously good search trump's the dark mode tho.
- irrational 4y agoI use duck duck go. Recently someone showed me their screen where they were using google to do a search. I was absolutely aghast. The last time I used google when you searched for something you saw a simple text list of sites (which is how DDG still works). Instead the google results were… a disaster. You had to scroll through some much garbage before finding actual search results - a list of sites. It was like google was saying, “here, look at all this trash instead of clicking a link and going to a different site”. When did google become so bad?
- 2Gkashmiri 4y agoads. SEO. i have a guy in digital marketing tell me that his friend does SEO and he does wonders with obscure keywords and shit. that friend is a freelancer and earns a good payday. When you want to insert your brand in every fucking imaginative keyword as opposed to people "searching for something", why does internet advertising revolve around everyone assuming every person googling something "WANTS TO BUY SOMETHING"?
- charcircuit 4y ago>why does internet advertising revolve around everyone assuming every person googling something "WANTS TO BUY SOMETHING"? Because people googling SOMETHING are more likely to buy SOMETHING than people googling SOMETHING_ELSE.
- RileyJames 4y agoIf every heavily SEO’d result was produced by a company that produced a directly relevant product, I don’t think we’d be as disappointed with the content. The truely garbage content is produced as cheaply as possible (scraped, generated from a data source or generated via “ai”) to capture advertising revenue, often via sub prime advertising networks (or a number of middleman networks). But to your point, not everyone wants to buy something, and not everyone needs to. Much of the content out there is simply trying to capture your attention and make you available to some of the worst advertising and ad networks (read scams, lead gen, fake buttons, affiliate crap).
- 4y ago
- nix23 4y agoYaCy?
- y42 4y agoI totally like the idea but I dare to doubt that this would solve the SEO problem. Website owners who are participating in those notorious affiliate programs or earn money with ads will still use the search engine to drag people onto their generic websites, using methodolgies to fit the search engines ranking mechanisms, no matter if they are public or not. SEO and all it's results seem to be immanent to the system.
- daoudc 4y agoI'm hoping that having community moderated search results will limit this problem. The problem will be SEO people trying to infiltrate the community, which may be tricky to solve. But Wikipedia has had to grapple with similar problems, so I think it is solvable.
- 6510 4y agoSell weekly SEO guids to rank well and punish last weeks tips. (joking) What you want to do is put the SEO monster in front of your cart and make it do useful work. You basically got an army of hard working people with money to burn who will do anything. What is there to complaint about?
- gnramires 4y agoYou can have a semi-manual system, with user input. So websites filled with SEO get manually flagged by users. Of course, you need to trust your users as well... this can be done with a Web of Trust-style reputation system: users endorse each other, and you can build a reputation graph/tree that traces reputation and easily cull bad subtrees. If this endorsement system conserves reputation (you give a fraction of your reputation away each endorsement, and new reputation is never created), then it becomes sybil-proof, where it's not advantageous to create say millions of users to increase reputation. (previous suggestion here: https://news.ycombinator.com/item?id=31585340 https://news.ycombinator.com/item?id=31585340)
- pronlover723 4y agowhy can't all the users that are being paid to promote spam endorse each other? And won't distributed trust systems just make some people "trust billionaires" and others "trust impoverished" though no fault of their own? If everyone trusts Oprah she'll have the types of same power people complain about billionaires people now. Basically influence. Also, anyone who's close to her gets the blessing of her influence where as if you're far removed you get nothing. Seems like the just reproducing the exact thing so many are trying to counter.
- dmje 4y agoHe's doing that thing [1] where's he's writing about a thing and presumably wants me - the interested reader - to know more about that thing because it's the thing he's spending all his time on, but he gives zero navigational options to his thing. So as that interested reader, it's down to me to find the name of his thing [Mwmbl] and then (hilariously, given the context!) use a search engine (probably The Evil Google) to find HIS thing. Seriously, people, if you're writing about anything at all, making assumptions is always a bad idea. If you're writing about a product, make it more than easy to get to it. Provide plentiful CTA's (that's Calls To Action, defined so as not to make the same mistake of assumption) - links, bittons, a big banner at the top: ("I'm building a non profit search engine called Mwmbl! Find out more"). K, thanks, </ moan > [1] https://news.ycombinator.com/item?id=31494925 https://news.ycombinator.com/item?id=31494925
- elashri 4y agoThe first sentence contains the link to the project.
- daoudc 4y agoThat word (Mwmbl) is actually a link, but I think the link colour is not different enough from the text. Let me see what I can do.
- dmje 4y agoIMO you also need an about page. Why are you doing this? Who are you? Why should the world pay attention? And stick them in your navigation. Overstate your cause rather than assume that people will get it. It's a rad product you're making! Don't undersell yourself! :-)
- 6510 4y agoI use to run a script on a blog that replaced a few thousand words and word combinations with links to the evil that shall not be named.
- lukeschwartz 4y agoNever thought real-time searching will be so cool https://mwmbl.org https://mwmbl.org
- cgrealy 4y agoThat’s a terrible name
- nitred 4y agoI disliked it too intially but it had grown on me after a couple of days. As to why the name was chosen and how to pronounce it are in the FAQ [1] [1] https://github.com/mwmbl/mwmbl#how-do-you-pronounce-mwmbl https://github.com/mwmbl/mwmbl#how-do-you-pronounce-mwmbl
- TulliusCicero 4y agoThe very fact that people look at it and think "what the fuck is this" is a problem, even if there are decent reasons under the hood somewhere.
- labster 4y agongl rand 5la op url is confusing af, not lgtm.
- 6510 4y ago> There are two ways forward that I can see: > The paid subscription model > Donation funded, non-profit model No! There is a 3rd! You could do a search app eco system where you leave the unlimited overly complicated puzzles a search engine could address as an exercise for the user. I always have a bazillion ideas but couldn't think of a single good phone app before mobile phones. I mean, should I want my phone to be a gaming console? It seems ridiculous. Writing is writing books, all other kinds are watered down. Do I want to write books with an onscreen keyboard? It all sounded idiotic, nothing worth using. But the idea you mention, typing an overly popular domain name without extension should take you to the website directly... What you are trying to say IMHO is CLI! Search is just the failback if the provided query/instruction doesn't make sense to any of the apps. I cant think of many but there are no doubt thousands of activities that could benefit from an at least somewhat themed search engine. An app could be a biochemistry web directory that ranks results from a chosen sub folder above the normal results. Any FOSS or other company could create a web dir tree with the few or many pages about it self. A check box lets you pick the ones you want to query. Normal results go under those results. The biochem wont bother you when searching for pokemon. People love my stores. What they really want is to see illustrated results from my inventory above all other results. Uncheck the box if you are not in the mood. (edit: I'm joking of course but I do have a good fews shopping apps that I actually use)
- pasdechance 4y agoI go Mojeek, then DDG. I kept forgetting about Kagi. I have a login for that. Yep.com has a different model, I haven't read into it far enough to decide if they actually do as they say they'll do.
- soperj 4y agoI really like your idea to have users help you crawl the web. I just don't like what your extension is asking for when I try to install it: - Access your data for all websites - Monitor extension usage and manage themes
- daoudc 4y agoFeel free to read the source code if you're concerned about what it does!
- tyropita 4y agoQuite a neat way to crawl websites using a browser extension. That by itself is a form of donation to the search engine. Maybe in the future you can have dedicated software for self-hosted clients that users can run to crawl and index websites for mwmbl? Kinda like folding@home. How are the batches of URLs to be crawled generated/discovered and posted at your API? How do you deal with duplicate crawls?
- thatwasunusual 4y agoI have also thought that distributed crawling with the help of browser extensions, and/or clients like folding@home, could be a good idea. But how to deal with "spam injections"?
- 6510 4y agoYaCy just crawls the results [again] locally before showing them to you.
- thatwasunusual 4y agoMaybe I misunderstand, but doesn't that mean you lose the benefit of having distributed crawlers if everything has to be crawled (again) locally somewhere?
- nix23 4y agoYaCy can do distributed crawling and exchange the Indexes (in Peer to Peer mode). I have some node's who just receives and send indexes without crawling (much less storage intensive).
- tyropita 4y agoAfter a certain scale I think you can let clients do double-work and let the most common crawl data, among different clients, win. And since you control what URLs need to be crawled, you protect yourself against rogue clients sending arbitrary URLs. There certainly are a lot of elegant ways to reduce spam for this particular problem imo.
- eterevsky 4y ago> Google tries to work out which sites are interesting by how long you spend on the site. How would Google know how long you spend on the site? It only sees what links you clicked and doesn't know what happens next. (Unless the website uses Analytics, but Analytics doesn't affect search ranking.)
- IshKebab 4y agoThey know if you click a result then 10 seconds later click a different result.
- defrost 4y agoLet's not forget they also have a very popular browser that itself collects and bundles back home a lot of usage infomation from most of its user base: https://www.ghacks.net/2021/03/16/wonder-about-the-data-google-collects-in-chrome-and-links-to-you-now-we-know/ https://www.ghacks.net/2021/03/16/wonder-about-the-data-goog...
- eterevsky 4y agoI think this is a bit misleading. Chrome gets access to various data to enable various features -- for example to pass your location to the website (when you allow it to do so). Browsing history is used e.g. to power URL auto-completion. This doesn't mean that this data can be used to inform Google Search ranking. That would be very shady and potentially illegal. I work for Google, and even though I do not work specifically on Search ranking, this doesn't sound like something that could be happening.
- defrost 4y agoI'd be surprised if they don't have the capability to enable "linger time" statistics that collate frequency of new web site loads, memory demands, etc for performance. This relates to the "how can they know how long I look at a web site for" asked above - if not specifically they do at least know the answer stochastically.
- mattrick 4y agoOne feature that I really wish more search engines would have is the ability to blocklist certain domains, particularly ones whose results are never relevant or helpful to the query itself (Pinterest, Quora, etc). It could even be used as a factor in the site’s search rankings.
- norman784 4y agoI think kagi does that, I use while was in beta, you also can assign a priority to sites, like normal or boost. Kagi is a paid service and doesn’t shows you ads.
- andrelaszlo 4y agoYeah, each user can apply their own weight per domain: block, lower, normal, higher, pin. It's really useful!
- kevincox 4y agoIt does. This is one of the best features.
- dazc 4y ago> It could even be used as a factor in the site’s search rankings. They did that for a while without considering the obvious downside - the incentive to mass blacklist your competitors.
- klabb3 4y agoBack in the day it was just w3schools i wished I could block. I still do, but now there are countless others as well. Quora in particular.
- cfiggers 4y agoHave a look at Brave Goggles: https://search.brave.com/goggles https://search.brave.com/goggles
- throwaway2056 4y agoI bet two-thirds of hn crowd is some how affiliated to the progress of search, ads, lead-generation, analytics, user-tracking (FAANG) etc. Think of their children...
- benrmatthews 4y agoHow is the name of the site pronounced? “Mah-wim-ball”? Google’s name was so ubiquitous it became a verb, Duck Duck Go is a smart memorable name. Mwmbl is a challenging product name, even if the .org domain name was available.
- lukeschwartz 4y agohttps://github.com/mwmbl/mwmbl#frequently-asked-question https://github.com/mwmbl/mwmbl#frequently-asked-question
- mark_and_sweep 4y agoWhy not call it Mumble then? I was also confused by the name. That's not necessarily a good first impression.
- muthuh 4y agoThis particular letter-sequence is an extremely unfortunate choice and should be reconsidered imho.
- noisenotsignal 4y agoAs the underlying project discussed by this post is a search engine, I searched for “mwmbl” on mwmbl.org [0], and no results were found! Relevant results like the main site and GitHub repo show up when searched on Google or Kagi. [0] - https://mwmbl.org/?q=mwmbl https://mwmbl.org/?q=mwmbl
- CobrastanJorji 4y agoI searched for "Google" on mwmbl, and while the first page of results found many results, including Google Patents, Google Bug Hunters, and a Google Books page on a 2004 book about Anarchism, it did not find Google's home page. I searched for "Elephant," and I got a Wikipedia page about a specific Elephant statue at Coney Island, a UK elephant charity, and a blog post about Haskell ("the elephant in the room"). It's unfair to poke fun at a very small project that admits that it is far from done yet, but it's gotta figure out a way to crack the "which pages are most likely to be relevant" problem or else it's not going to be useful.
- noisenotsignal 4y agoTo be honest, I didn’t entirely mean it as a comment on the search quality. But you would expect a search engine to return results about itself, and it was amusing it didn’t!
- daoudc 4y agoYeah the search quality has a way to go!
- guerrilla 4y agoPlease add it to [1] since Firefox (absurdly!) doesn't seem to let us add arbitrary search engines anymore. 1. https://addons.mozilla.org/en-US/firefox/extensions/category/search-tools/ https://addons.mozilla.org/en-US/firefox/extensions/category...
- tgv 4y agoIt does, I just tried it with 102.1. You can add a "keyword" to use a search engine, and once you've done that, you can add it as a search engine to the url+search bar (the main one on top, I don't know how it's actually called) as well, and you can set it as your default search engine.
- guerrilla 4y agoI mean I used to have that but I seem to now. Did it move? This is 102.0. Edit: You can make a bookmark and add a keyword, but that doesn't help me use it as a search engine. It used to be you could just create one from the contextual menu, what is this convoluted process.
- tgv 4y agoThe steps I took were: 1. Go to kagi.com 2. Right-click on the search field and choose "Add a keyword for this search..." 3. Fill in the keyword (e.g. kagi, as I did) 4. Right click in the top url+search field, and choose "Add Kagi search" Now you can search via keyword, you can change to kagi while typing, and you can set the default search engine to kagi in the preferences.
- happymellon 4y agoFirst thing I try is: Star trek imdb First result is startrek.com, second result is Star Trek into Darkness IMDb but 3rd is xkcd. It then goes off into Q and William Shatner Wikipedia links and Muppet Movie IMDB in Russian. I tried putting a plus in front of IMDB and quoting Star Trek. It doesn't seem to be able to find Star Trek on IMDB. I admire the concept, and it is extremely fast.
- daoudc 4y agoOur index is still very small. Help make it bigger!
- happymellon 4y agoSorry, I can see that sounded bad. I just meant that it sounds like we have a ways to go. I will help!
- daoudc 4y agoNot at all, feedback is very welcome!
- happymellon 4y agoI wouldn't be surprised if you don't want to talk about it because not could try and avoid the pitfalls, but what is your plan for avoiding bots trying to overwhelm the search scraper bot? Are you looking to build a trusted network who can verify and validate other users responses on an undefined period?
- TulliusCicero 4y ago> Instead of looking at how long people spend on a site, we would encourage users to give explicit feedback on rankings and use this to improve our ranking system. While they're not wrong about how the way Google determines ranking has its issues, this way has its own set of problems. If you explicitly use user ratings as part of your rankings in some way, people can punish sites they don't like, ala review bombing on Yelp, Steam, etc. Not saying it's necessarily a bad idea because of that, but I hope they don't fall victim to the mentality of, "let's just trust the users" as an ironclad rule, because that doesn't always work out well.
- whoknows2 4y agoAlso let's not forget Reddit. There's a lot of bullshit on top.
- zo1 4y agoI went back to Reddit a week back after a 1 or so year, and what I saw was just disgusting on the main feed. One click, a bit of scroll and I was watching people calling republicans terrorists and being praised/upvoted/gifted. It's not just about bullshit at the top, it's also about echo chambers and the online equivalent of mob-behavior.
- lotsofpulp 4y ago> it's also about echo chambers and the online equivalent of mob-behavior. This particular example might be about the Republican party-wide support of echo chambers and offline mob behavior that led to an invasion of the building containing politicians certifying the vote of a newly elected leader and fueling disinformation about elections to weaken the trust and integrity of the system?
- AussieWog93 4y agoRegardless of this particular example, there has always been a pretty strong bias towards one side of US (and Australian, for that matter) partisan politics on the front page of Reddit; with over-the-top accusations towards the Republican Party as well as the LNP both receiving thousands of upvotes despite their poor quality.
- ilaksh 4y agoWe don't need another centralized search service. We need protocols for publishing and finding information that do not rely on servers.
- closedloop129 4y agoCan you elaborate? https://yacy.net/ https://yacy.net/ is a decentralized search engine. Are there shortcomings to the protocol?
- ilaksh 4y agoYacy is exactly the type of thing I am talking about. When I tried it, did not seem perfect but definitely a better overall paradigm. It's weird that more people don't know about it. But also there should be more than one decentralized search engine.
- siquick 4y agoBeen using Brave search for around 6 months and probably 10% of searches need the !g parameter added. Braves working well.
- pronlover723 4y agoThis seems really really naive. Do you really think a non-profit is going to fight the hordes of spammers, scammers, seo masses, mechanical turk hordes, etc that are going to game your system?
- daoudc 4y agoWikipedia has to deal with similar issues. It's not easy, but it should be possible.
- carschno 4y agoNot sure I missed the sarcasm here. Why would a non-profit not fight spam etc.? Why would a for-profit fight spam do so? I see the largest internet companies, including Facebook, Twitter, and also Google, fight spam and other harmful content only to the degree absolutely necessary to stay somewhat usable. Which makes sense because it's costly and does not generate profit. I would expect a non-profit, however, to focus much more on fighting harmful content because it centers around the user experience, hence quality of the content. I don't see a guarantee this works in practice, but the respective incentives seem clear.
- joshuamorton 4y agoThe general answer here is that a nonprofit cannot maintain the (massive) amount of resources it takes to address spam/abuse/blackhat SEO etc, while a for-profit entity ostensibly has a profit motive to do a decent job, and the resources to do so. When a for profit entity is more successful at fighting abuse, their users are happier and they sell more ads and so can devote more resources to fighting spam. When a nonprofit successfully fights spam, they don't get more resources, and the spammers upgrade their toolboxes, because they do have a profit incentive.
- closedloop129 4y agoLike @daoudc writes, Wikipedia shows that the nonprofit can maintain the resources because they attract volunteers. If you create a search engine where users can report spam and get some form of karma for valid reports that is shown in their social network, then it's quite likely that the users have enough momentum to get ahead of the spam.
- labrador 4y agoI don't know what the world needs but I need a personalized search engine. I would like to filter out anything to do with sports. I would like filter articles that contain marketing jargon and technobabble. I would like to filter articles written below high school grade level. And so on.
- rambambram 4y ago> I would like to filter out anything to do with sports. Amen. How many times I searched for an astronaut's name - or any other person with significance - and instead I got results for some pro sports guy. I guess 'significance' is subjective and I should search more specific.
- guerby 4y agoI would love to have a search engine with buy/nobuy tag: shows only shops with "buy", and no shops with nobuy in the search results.
- labster 4y agoSo “nobuy” limits results to astroturfed product reviews and link farms, then?
- unsignednoop 4y agohttps://en.wikipedia.org/wiki/Quaero https://en.wikipedia.org/wiki/Quaero
- iamjbn 4y agoIf you need a good search engine, pay for it. Business needs to make money, it’s that simple.
- avgcorrection 4y agoThis search engine is supported by The Bill (and Melinda?) Gates Foundation, The Organization for Promotion of Democracy, The Organization for Prosperity, The Organization For Truth And Transparency And Against Fake News, The Organization Against Renegade Knowledge, The Organization For Helping Silly Citizens Think Better, The Organization For The Truth About Qatar, The Organization For Freedom And Good Things And Not At All Tied to the CIA, and some other folks.
- skitout 4y agoAbout the funding model, it would be great if normal donation works. If not, I think the way some WEB3 projects are funded maybe an interesting inspiration (not talking about Ponzi scheme here). Many projects are "non profit" and sale tokens before the service is 100% ready. It funds the amelioration and scaling of the project. And the possibility to resale the tokens at a higher price in the future sometime attract token holders and often increase the "motivation" of the token holders / supporter of the project... fuel the community (money is only part of the motivation). Here token could be associated to symbolic "privileges" (badge, access to early releases), or governance (taking part of some votes). This system have clearly some drawbacks, but allows sometimes to increase the number of early users and supporters, and get more funding while staying a non-profit.
- kolinko 4y agoPerhaps a better approach would be building an open source www index or even a full current cache - as an enabler for people to build their own search engines? Right now it is extremely difficult to build your own web crawler that would compete with Google. And that is not because of the technology, but because multiple sites will prevent your bot from accessing them if you're not Google or Bing - either through robots.txt, or through directly banning your IP if it's trying to crawl and it's not a confirmed google-bot. Having a non-profit, open source, crawler that keeps an up to date index (or web cache) of the web would help competition spring up.
- nix23 4y ago> Perhaps a better approach would be building an open source www index or even a full current cache - as an enabler for people to build their own search engines? That's a excellent idea! In the spirit of open-data, and people can do with it what they want.
- artificial 4y agoI think this is a great idea. How does this work with copyright? Search engines seem to be able to download a reproduce content from scraped pages (and wrap it in ads, and derive content from it) this is called “indexing” when they do it but scraping when everyone else does it.
- bosie 4y ago> and wrap it in ads, and derive content from it i am probably missing something but can you give an example where this happens?
- Flashtoo 4y agoE.g. on Google if you search for "how to tie a tie", a little info box may pop up with step by step instructions. This content is taken from some website, but that website gets no page hits or ad revenue. Instead, Google gets to serve ads on the search engine results page. (I don't know if this happens for this specific example, but Google does this for some searches)
- aaron695 4y ago
- golf_mike 4y agoHope this takes off. Also hope he works on his math when it comes to funding, 1% of 40 billion is 400 million, not 4 :p
- daoudc 4y ago1% of 1%
- infinityio 4y ago> Just 1% of 1% of this His maths is correct, just an easy-to-misread phrase
- golf_mike 4y agoI hope my reading will improve then!
- jsmith99 4y agoI think Google's results could be a lot better but I'm relatively ok with my search being provided by a for profit company. Their incentive is to get me to want to use their product. A non profit with that much power might be more tempted to manipulate search in ways that suit their personal preferences.
- klabb3 4y agoI don't mind profits but monopolies are always bad for the consumer. I wonder what innovation would have happened in search if it had been more competitive.
- mordae 4y agoI think that we are past the point the search engine could just crawl the web and rank results based on some heuristics. We need both community curation and get librarians involved with their classification systems, because in 3 years the results are going to be dominated by automated GPT-xy content farms. Case in point: www.forkandspoonkitchen.org The first search engine that provides community curation and manages to get most tech-savvy people on board, classifying the content for free, is going to reign in the upcoming decade as Google loses its grip.
- M0r13n 4y agoI am conflicted when it comes to stuff like that: - on the one hand I really want free, open and non profit services to succeed - at the same time I greatly value the user experience Don't get me wrong: These two things can go hand in hand. There are tons of good examples out there. But, the closer you get to classic user-centric applications and leave the software developer bubble, the greater the discrepancy becomes in my experience. Brave, DuckDuckGo, Firefox and so on are desirable. But I always feel like I am missing out on the UX. Google still yields better search results FOR ME(even with all those ads and clickbait). Firefox still feels a bit dated and slow compared to Chrome. I value the positive effects of free software so much that I am willing to accept limitations in usability in the hope that it will improve over time. But I feel like it should not be this way. I can't support every project financially or contribute to its success as a contributor. My time and financial resources are limited. I haven't really found a solution for this problem. My best guess is that the government should intervene in the free market and install market barriers to tame giants like Google. But this is repugnant to the liberal in me.
- renonce 4y agoThe biggest challenge with making a search engine is to combat adversarial SEO. It's an issue that's very easy to be overlooked when you are small, but at Google scale, your enemies have billions of dollars to make from your visitors. I bet Google spends at least as much to combat that, and it's extremely hard to deal with while being open-source. It's useless to call for a non-profit search engine without tackling this very core issue.
- solarkraft 4y agoAre your sure they're combating it? It seems like they've given up.
- renonce 4y agoYeah... but if Google is giving up, who would ever be able to do it?
- solarkraft 4y agoI think it's possible - red flags (for example blog spam or commercial sites) seem easy enough to catch for a human; probably in an automatic way too, at least as long as you're small enough that they don't specifically target you. Google's search results are so bad I can't really alledge incompetence here but have to wonder whether there's some different motivation. Maybe it's that low quality search results tend to be plastered with ads, which they get a cut on.
- Nextgrid 4y agoThey haven't given up, they are staying exactly where they need to be. All the mainstream search engines' priority is to maximize ad clicks/impressions (or collect data to target future impressions), either directly on their own property, or indirectly when linking to websites that embed their ads. There's no reason why they can't detect ads or analytics and use that as a negative ranking factor (so that all other factors being equal, a non-ad-infested result would rank higher than the ad-infested one), but this would go contrary to their business model.
- solarkraft 4y agoThe internet, somewhat ironically, really needs a search engine that works in the current day. You can't find anything anymore. It's like Google has been un-invented. Hopefully some day soon the internet will be searchable again. Thanks to everyone involved in attempting to make this happen (preferably in a non-profit-maximized way). (Said before at https://news.ycombinator.com/item?id=32034390 https://news.ycombinator.com/item?id=32034390)
- dageshi 4y agoI have a feeling that people stopped writing actual good website content for certain topics and moved to places like youtube instead. Youtube frankly offers better monetisation and most importantly easier user retention via subscriptions.
- solarkraft 4y agoWhile monetization is certainly a great thing for a group of people wanting to go after things they and others enjoy, it's also exactly what incentivizes people to game the system and pretty much what brought us into our current place of spam, spam more spam and barely any way to discover things that are actually good. I tend to think the quality of the content tends to be better when people don't think about stuff like user retention or subscriptions, but rather how it will actually reach people that care. Good search/curation is a key component for that. Of course such a world free of implicit monetization will require it to be explicit (Patreon-style), but that should massively realign incentives.
- dgudkov 4y agoIt's unfortunate that with all the immense value that search engines provide the idea of paying a small monthly or annual fee to use a search engine is incomprehensible for most people.
- AvSaba 4y agoA small monthly fee could mean a lot of money in the developing countries, plus there are countries where most people don't have access to international banking system. Such a system only makes access to information for a lot of people almost impossible.
- txtsd 4y agoFor the price Kagi asks, I could buy 2 weeks worth of groceries.
- gfxgirl 4y agohttps://neeva.com https://neeva.com is trying
- HWR_14 4y agoA small fee to a SV resident is a larger fee to most of the people in the rest of the US/EU is a prohibitive fee to many others.
- voltagex_ 4y agoThere are so so many anti-scraping sites around - it's very difficult to do without pretending to be Googlebot or whatever.
- epolanski 4y agoThe way I see it, Google is no longer in the business of searching websites, but in the business of ranking them from at least a decade. I still remember helping a friend finding informations on the accounting balance of Rome's, Italy, public transport, and finding the most relevant link buried deep at around page 20. The first 15 pages were almost completely news websites with completely irrelevant news to the search query but they would consistently rank much higher.
- mcv 4y agoI've been thinking in this same direction. Especially the community-driven part. Google seems to be more interested in what corporations and advertisers want, rather than what users want. With their tendency to crowd-source their AI training, I'm surprised they don't let users vote on search results. If I were to make a search engine, I'd definitely give users more control over their results. Block crap sites, vote up your favourite sites, vote down questionable sites, maybe different context profiles, because if you're searching for Java in the context of vacation or news events you want different results than if you're searching for it in a programming context. There's so much that search could do better than what Google is doing, but I'm not doing it because it's way too much work, and it requires serious resources to index everything.
- cvccvroomvroom 4y agoThe world needs a non-profit and co-op social enterprise most things.
- frozencell 4y ago> Just 1% of 1% of this would be more money than I’d know what to do with ($4m). Not knowing what to do with $4m means the failure of the education systems.
- gattilorenz 4y agoTaking everything literally, this comment included, is the failure of the education systems.
- dalbasal 4y agoThat the world needs a non-profit search engine is near trivially true at this point. So good luck Daud. I think the pertinent question though, is what's the best way to demonopolize search. Maybe the answer to that is non profit, maybe something else. Google has a most search users. They have an even higher (much higher) portion of search revenue and essentially all of the sector's profits. One advantage a non profit might have is going after the low profit parts of search. Use cases where Google is likely to be under-serving users. Also, search isn't just websearch anymore. It's a way of calling a calculator, translating, etc. It's a text box that does stuff. The newest gen of language models may be the technical catalyst for some rapid evolution in the "clever text box" space. Google is obviously super active in this space, but shifts are a good time to get in. Where would you skate, if you were skating towards where the search puck is going?
- tommoor 4y agoI wish the Google One subscription would just remove ads from results.
- amadeuspagel 4y ago> Fast ... Instant Search Great idea, and "instant search the web" would probably a better pitch then "non-profit search engine". Interesting argument that google doesn't do this because it isn't compatible with their ad model, but that doesn't mean a new ad-funded search engine can't do this. For google it might be billions of dollars in lost revenue while they adjust their ad model, a new ad-funded search engine wouldn't have this problem. > Frictionless ... For example if you are typing “facebook” or “hmrc login” you could go straight there from the address bar. No thanks. I sometimes do search for "company name" looking for the wikipedia article for the company, or news about the company, or information about the company in general. If you used facebook before, then it's going to autocomplete as soon as you type "face" in your addressbar, and you won't need the search engine. So if someone searches for facebook, they're either using the browser for the first time, or they're looking for information about facebook. Latter seems more likely.
- larsrc 4y agoDisclaimer: I work for Google, though far away from Search. Regardless of search engine design, there's HUGE money in SEO. Any successful search engine will be gamed. Do you have the developer power to go red-queen against all the large companies in the world?
- DisjointedHunt 4y agoAnd that money goes to support many industries beyond search itself. The author really needs to get off a computer for a minute and understand the economics of the web as it stands today and how the free ad supported model supports millions of peoples livelihoods before jumping to "This is all bullshit"
- vineyardmike 4y agoWhile I agree that they may need to better consider the economics - both of the engine and the websites that may SEO it- that doesn’t mean we should just assume as supported is the way to go or the only way to support people. The economy of today looks different than 20 years ago or 20 years before that. Doesn’t mean we shouldn’t grow and change.
- DisjointedHunt 4y agoI respect that. I guess my point is "Grow and change with a decent understanding of what the present state enables" A search engine like Google isn't just a search engine as the author describes it. It is a very integral part of the economy of the internet and just labelling a simplistic interpretation of the present state as "Evil" with an academically poor write up of what a viable alternative is does little good.
- 6510 4y agoI want to write articles and read articles written for me by others. Ideally as few as possible should profit from this process. As google is now a turd, not just no longer capable of delivering this service but actively destroying the good part of the web by refusing to index it. It is my attention, it doesn't belong to anyone else. My access to information and educated opinion is a far more integral part of the greater economy. Google is like a screaming man at a town meeting making sure no one else can get a word in. The meeting is now pointless.
- phtrivier 4y agoThe "funding options" part has the unsurprising blind spot that, maybe, a search engine is the kind of basic infrastructure that ought to be paid (at least in part) by... The taxpayers ? I have a long standing bet that, at some point, some company will be "globalized" (operated under some common funding by many different countries, like many research projects or defense organization or aid funds, etc...), and the "search engine" part of google is the prime candidate. That being said, I'm from Europe, so "sharing the cost of something useful" is not culturally untolerable. Far fetched and controversial opinion, I know. We'll see.
- Timwi 4y agoThe author is not going to be able to convince governments to fund it unless they can first demonstrate a viable service.
- vineyardmike 4y agoConsidering their goal was to reach £50 a month to upgrade their server… I think any hope at convincing major governments is a few years off. But I do wonder if the national archives or the library of congress could be a good host for this sort of project. Not sure I agree it should be run by a government… most don’t have great histories when put in charge of gatekeeping access to information.
- ZeroGravitas 4y agoI'd quite like a browser extension that records all my searches and where I end up, just for my own review. I feel like many of my searches aren't actually searches, but I can't quantify that at the moment. Feels like that would be good info to share, once it's depersonalised.
- DisjointedHunt 4y agoThis is the kind of idiocy that makes me despise the developer community every time i see something like this. It is one of my pet peeves, so if you're going to have an opinion on this comment, please, read the whole thing. The ad supported free internet is one of the most important business models the world has arguably ever seen. Very few can argue with the fact that poor kids in developing countries over the past two decades and longer have had their lives changed beyond anyones wildest dreams thanks to the free resources at the tip of their fingertips. On the same note, much of the wealth accumulation in the developer community has been on the backs of this very business model. The immense demand for dev talent and the astronomical salaries paid out is a consequence of the difficult financial choices made by so many before us. When i read absolutely low-effort activism such as the text in the link about how(paraphrasing) 'sEaRcH eNgEnEs mAkE mOnEyY" and thus they are bad. I'm astounded at how intelligent people who can write code can simultaneously be so fucking moronic in their grasp of economics. The web is an ecosystem. There are always going to be incentives that don't fit your moral compass that are getting optimized for and against. The answer isn't to burn it all down and shit all over a business model because it apparently doesn't fit your childish understanding of the ideal. By all means, compete, but atleast try to understand the various actors and participants in this complex web of entities and what role they're playing in the flow of investment, content, data and economic activity that is far more nuanced than "wEb rEsUlTs wIll B beTtTeR iF nOT oPtImiZeD fUr $$$ " Face fucking palm
- klibertp 4y ago> I'm astounded at how intelligent people who can write code can simultaneously be so fucking moronic in their grasp of economics. Absolutely nothing surprising about that. Intelligence without knowledge is not very helpful. If all you know is how to write code, you will suck with other things, even if you're intelligent. The bigger problem is that people tend to downplay the knowledge that's required to do something, simply because they do not know how much they don't know. It gets worse the more intelligent you are because you're more confident in yourself then. Case in point: name of this project. I've read it like 10 times on this page already yet I still can't spell it from memory. I could paraphrase: > I'm astounded at how intelligent people who can write code can simultaneously be so fucking moronic in their grasp of marketing. (but won't, since, as I said, it's not astounding at all; just for illustration purposes)
- Timwi 4y agoLove the idea and the project! However, if you are aiming to become popular, definitely the first thing you need is a better name than “Mwmbl”.
- kebman 4y agoIt's perhaps a bit on the side but still part of the topic of search. Have you noticed how newspapers systematically do not supply a clear source for their articles? It's especially prevalent on political cases where there are easy-to-link paper trails. This makes it a lot harder to find the source for their article, so you end up just taking their word for their angle on the story. A great recent example is Biden's Executive Order on the protection of women. When the newspapers writes about his EO, they're never doing it form a neutral standpoint. In this case they're either pro or anti abortion. But if you want to know the contents of Biden's EO for yourself, then you're forced to search for it. And depending on the search engine, that might also be hard because also search engines are politically biased. Just so we're clear, this post isn't pro or anti abortion. Instead it's an example on how newspapers systematically force you to take their word for their angle on any given news story. So if you want to know source material, then you're forced to search for it. And when you do search for it, you're then at the mercy of the political bias of the search engine. For that reason I'm not so sure a non-profit search engine will make political biases go away, especially when you consider what happened to Wikipedia. While not a search engine, it is a non-profit and communal project that set out with the ideal of being truly neutral, but in the end it failed at that, and some would say spectacularly. And the main reason is exactly bullshit, or rather the BS that comes with political bias. Don't get me wrong, it's still a great source for information, but when you search for any topic that is in any shape or form politically sensitive, then you have to know about Wikipedia's clear political bias beforehand, or else you might take their angle as gospel. This is especially insidious when it comes to search engines and also social networks, because most people assume that what is shown to them there is neutral, or at least coming from a friendly party. But then it turns out, that's not always the case. When you systematically get biased information, then it's a democratic problem, because it prevents people from making up their own mind about political topics. Thus when people finally vote, the risk is that we get a society that does not reflect peoples actual opinions. I think most people in here has been on the receiving end of that, no matter which side of the aisle you're on. And the result is always resentment and bitterness which in turn does not make for a healthy democratic environment. Instead the political bias should be more clearly visible and out in the open on both newspapers, encyclopaedias and search engines alike. And while a non-profit search engine would certainly save you from corporate interests, it still won't save you from political ones, though it might be a good trade-off to save privacy.
- Sujeto 4y agoMost of the time you want answers from a certain place, be it reddit, or stackoverflow. It's usually easy enough to add the site, by simply writing its name, but it could be easier. I'm thinking right now this: - Compile a browser without the cross origin limitations - Make a site that uses iframes with all the answer-providing websites in it - Simply focus the text input in the site/iframe you want and search away - Have a way to open the results in your main browser or just use that patched browser Like one of those internet explorer toolbars, except they cover the whole area.
- 8organicbits 4y agoRecompiling a browser is a pretty heavy lift. DuckDuckGo allows one text box to be rerouted to a different search engine with a "bang". https://duckduckgo.com/bang https://duckduckgo.com/bang You could also build your own, most search engines have a specific pattern to how they encode the search term in the URL. Although, I suppose that doesn't support auto-complete
- closedloop129 4y agoWhy do you need to show the sides in iframes? You could submit your query to your target sites like this: https://searchaggregate.com/ https://searchaggregate.com/ [1] [1] https://news.ycombinator.com/item?id=31722417 https://news.ycombinator.com/item?id=31722417
- leobg 4y agoA recent comment here mentioned search in early browsers (1991ish). The browser would fetch all links from the current page n levels deep in the background and uses that to build a local index. I wonder if something like that could work today, only with the index being shared across the user base. The benefit would be that it’s a decentralized system. No giant infrastructure required which needs to be paid for by a big corporation. Basically, the infrastructure needs would be outsourced to millions of devices. And for websites, users and crawlers would be the same thing. Which is to say, you cannot block one without also blocking the other. It could also add feedback mechanisms. Active ones, such as commenting on pages and discussing them, as we do on HN. But also passive ones such as tracking how long the user interacted with the page, to score the value of pages/domains and improve the ranking algorithm.
- nonameking2026 4y agoBuild one. DAO could be a good soil for this.
- O__________O 4y agoAppears you’re in the UK, is that where you intend to registered the non-profit? If so, in the UK, what are the real costs of forming a non-profit, keeping records, generating reports, (shut it down), etc.?
- O__________O 4y agoWorth noting attempts at non-profit search engines are not new. In 2015, Wikimedia Foundation attempted to start one called the “Knowledge Engine” using at least $250,000 from a grant. Wikipedia likely started the project as a response to Google’s use of “knowledge panels” based on Wikipedia Creative Commons license alongside search results in 2012, which reduced traffic to Wikipedia. https://en.m.wikipedia.org/wiki/Knowledge_Engine_(Wikimedia_Foundation) https://en.m.wikipedia.org/wiki/Knowledge_Engine_(Wikimedia_... Also worth noting that Google is a significant donor (and now enterprise customer) of Wikipedia, but unclear if this had any impact of Wikipedia’s choice not to continue the project.
- daoudc 4y agoFascinating, thanks for this! Looks like it became controversial because of a lack of transparency.
- O__________O 4y agoWhich is the problem with non-profits, the transparency requirements are at best minimal and all non-profit vs for-profit means at a super high level is that there is: no equity distribution, government approves of its mission, and there is no distribution of excess cash flow. There frequently non-profits that use excess funds to unnecessarily expand beyond the original mission, for example Wikipedia — or that pay staff, especially executives way beyond what most donors realize. To me, being a non-profit is what it is, I don’t read too much into organization being a non-profit.
- jeffbee 4y agoI don't think this accurately represents Wikipedia's relationship with Google. Wikipedia is thrilled with the knowledge panels. It has dramatically cut Wikipedia's hosting costs and spread its reach. You're making it sound like Wikimedia and Google were antagonists over the knowledge panels, but as I understand things it was the opposite.
- laserbeam 4y ago> [google gets 40 billion a year from search.] I can’t even conceive how big it is. Just 1% of 1% of this would be more money than I’d know what to do with ($4m). Ouch. I wish you the best but that statement makes me lose hope. Employees are expensive. Servers aren't exactly cheap either. And unexpected mistakes along theyl way cost a lot.
- Pakdef 4y ago> Employees are expensive. Google has too many of them and it's probably why Google Search is not improving.
- Nextgrid 4y agoI wouldn't say Google Search isn't improving because of the number of employees. Google Search is exactly what it needs to be to serve Google's interests - it's just that their interests don't align with yours.
- samwillis 4y agoThe cost of building a “general” search engine for the “whole” web is astronomically high, in the 10s to 100s billions. It’s not achievable, Google were only able to do it by growing as a business at the same time as the internet itself. I don’t believe it’s possible to compete with Google (or Bing) by starting at zero. The route forward, and what should be advocated for, is a distributed network of search engines, each for a specific vertical. If it operated as a cooperative they could share expertise and technology, they could then build a “meta” search engine for the co-op that combined all the results from the specialist niches. Each member basically “owning” the “franchise” for a specific type of search or category. So, I don’t believe a single non-profit is the answer. More a co-op type arrangement where the co-op organisation (which may be a non-profit) has a mission to advance internet search through it’s network and strategic investment.
- geekamongus 4y ago> Google has an incentive to rank pages that contain Google ads because it makes them more revenue. Google has an incentive to rank profit-making sites higher so that they make more money. Is there evidence that they do this?
- Nextgrid 4y agoThey don't have to intentionally rank them higher, but this fact could prevent them from choosing to rank ad-filled sites lower even though it would do wonders for search result quality because ad quantity is usually a good proxy for spamminess and trash.
- javajosh 4y ago>Google makes $40bn...If I can create something that just a tiny fraction of people find useful, then I can create a huge amount of value. You conflate two meanings of value: monetary value, and intrinsic value. Search engines are intrinsically but not monetarily valuable to users. Search engines are monetarily, but not intrinsically, valuable to advertisers. You can get into trouble when you conflate these two meaning of "value". In fact, right here is the pivot on which the internet goes from an idealistic shang-ri-la for geeks, to a commercial hellscape for the unwashed masses. It is surprisingly easy to create intrinsic value with computers! You see it all day every day on HN: some geek had a thought, spends a weekend making it, and then deploys a solution. It is surprisingly hard to extract monetary value from an intrinsically valuable solution. In fact, I believe that creating artificial scarcity is the hardest part of building an internet business, requiring invention on par with the intrinsically valuable part - and yet its the very thing that idealists rail against. (And making something artificially scarce does seem morally repugnant. And yet I don't see any other way to pay developers. Full stop. Open source software + consulting fees is a good way to go, but that can't apply to hosted search for the public. Well I guess it could, you could teach businesses how to game your own engine!)
- amelius 4y ago> It is surprisingly hard to extract monetary value from an intrinsically valuable solution. This is like saying that our entire economic system is backward ...
- throwawaymaths 4y agoI have no idea what gp means by this. It can be surprisingly hard, but it's not always. Go buy a cooler of water bottles and sell them for 50c in the summer on the side of the street
- WJW 4y agoIt seems pretty obvious in an internet context. It's very difficult to make money with a product if the competition is giving their products away for free because it gets paid in a different way. Even your "sell water on a hot day" idea probably won't sell a lot if you set up shop right next to an enormous promo stand from a global bottled water company that gives away bottled water for free. (And due to the magic of the internet, every spot you pick is right next to a global competitor with deep pockets)
- mkozak 4y agoNot for profit? Ecosia is one. They do make money, but in general it's not-for-profit organization that use majority of money they make to plant trees.
- chris_f 4y agoWorth mentioning is the Alexandria.org project [0]. It is a non-profit search engine built on data from Common Crawl. The coverage is limited because of Common Crawl, but the relevance is decent. They also provide an API. I believe one of the biggest impacts toward breaking up Google's monopoly on search is making them open up access to their index, even requiring Google to provide direct API search access for others to build alternative search products. They have a search API today, but it is prohibitively expensive to build on ($5/1000 calls). I built a fairly popular search engine a couple years back, but the cost of Google's search API and increasing number of bot attacks make it difficult to reason keeping it online. [0] https://www.alexandria.org/ https://www.alexandria.org/
- bernardlunn 4y agoHas to be decentralized. Huge data centers need huge amounts of capital
- naillo 4y agoWith all the great progress in large language models lately, and them being excellent text compressors, I've started to wonder if you couldn't just replace a search engine with a like 100mb file full of weights that let you query essentially google scale results except all locally.
- FrenchDevRemote 4y agoYALM is 200GB and require 200GB of GPU Memory to run....
- naillo 4y agoYeah you picked the hugest SOTA one of them all, but there's smaller ones like this https://bellard.org/libnc/gpt2tc.html https://bellard.org/libnc/gpt2tc.html that run well even on CPUs that might run well fine tuned specifically on search results (or at least just code queries).
- FrenchDevRemote 4y agoThe only significant difference between those models is the amount of data, the main(or the 2nd most important) reason why you're using a search engine is how much data is available on it. If you want to search through an incredibly limited % of the web then yeah it can be a solution, but even the lamest search engine company out there would outperform a GPT-2 like model running from your laptop.
- nova22033 4y agoThis is a Elon "I'm going to build a hyperloop" Musk feel to it..
- deleted 4y ago[deleted]
- marginalia_nu 4y agoHonestly, the easy part is building a search engine, like just document retrieval stuff and domain ranking, SEO-mitigation etc. Anyone can build a Google '98 and get it to work well, not that hard, doesn't require all too much hardware. I have done that and got one running out of my living room. The tricky part, if you want people to use your search engine for more than the novelty factor, and what most Google competitors struggle with is drawing the rest of the damn owl. For example, commercial searches, local businesses, that sort of thing. As much as Google flounders with some queries, the overall package is still really good.
- albertopv 4y agoI agree, the whole package is the difference. I'm trying to use DDG as much as possible, but I live in Italy and for local stuff Google services are still unbeatable. E.g. Apple Maps used by DDG are a no go where a live for all but plain directions. OSM data used by many map apps is good, non commercial data sometimes even better than GMaps, but businesses data is only on GMaps, there's no way around that.
- jacooper 4y agoTry brave search, form my experience its really much better than DDG. And its independent.
- 1-6 4y agoAnd become the next Wikipedia? No thanks.
- Cupertino95014 4y agoTell me how this doesn't quickly devolve into a consensus-rules hellscape, where minority views are either ignored or certain minorities are artificially boosted. There is no way that design choices (especially the ordering of results) can be made in a way that pleases everyone. So either you dumb it down to the point of meaninglessness OR you enforce a mainstream-only ruleset.
- mrkramer 4y agoThe main enemy of a better search engine are casual users who are satisfied with Google's mediocrity and don't seek nothing more advanced and better. Power users are the one who suffer the most. Google will have to reinvent itself or it will eventually destroy itself with negligence of its core business. There isn't yet critical mass of casual users who think Google sucks, all they think is that the Google is internet. That's their intellectual level.
- bacan 4y agoThe issue with Google is SEO optimized spam sites beat out real content. Until we get rid of web advertising, it'll continue to be this way. Spam site operators have a huge incentive to get users to click on their links & provide them with ad-revenue For Google, they could make things so much better by down-ranking sites that show ads
- charlieyu1 4y agoWikipedia is non-profit and still manipulated by shills.
- formerkrogemp 4y agoI mean the IRS makes 990 forms publically available. They may be a year or two behind, but it's valuable financial and personnel data from nonprofits. EDIT: Ok, I see that this is about a search engine structured as a not-for-profit, not as a search engine for nonprofits.
- oidar 4y agoI get an incredible value from search engines. Google even (their shopping and book search features are very helpful). But right now, I am liking paid search as the way forward. Kagi is doing pretty good things right now. I love how I can up the weight of certain domains so that their results come in at the top without having to add site:awesomesite.com at the end of every search string. In fact, I can have 20 sites that I trust a lot that show up pinned at the top of the search for every query. It's 10 bucks a month, but I find it valuable.
- greenie_beans 4y agoDoes this work? https://projects.propublica.org/nonprofits/ https://projects.propublica.org/nonprofits/ America only
- Terry_Roll 4y agoWhen using a VPN to access Youtube, the adverts played to you will be in the local language of the VPN destination, yet Youtube can deliver the appropriate language content. Strange that!
- timbit42 4y agoHow many advertisers advertising in a VPN destination where a particular language is dominant would advertise in another language? YouTube doesn't care whether you can understand it. Those advertisers may not be looking to advertise to people who can't understand their language either. How many advertisers outside of that country are going to ask YouTube to play their ads in that country in a language different from the dominant language? Probably not many.
- factfindingisfn 4y agoIt definitly does
- mrtweetyhack 4y ago
- dredmorbius 4y agoFor an information based on standards --- HTML as a document markup language, HTTP as a transport layer, TLS/SSL for security, TCP/IP as an underlying networking protocol, among others --- one that is conspciuously missing is an indexing standard. That is, even if a site wanted to, there's no way for it to declare "I have content related to X". Even better would be if these indices could then be distributed in a cache-and-forward model similar to how DNS (another distributed discovery index) works. There was some exceedingly rudimentary attempt at this through elements such as keyword meta tags, but even at best these referenced a vanishingly small fraction of the actual content of a site or article. Sitemaps also address a component of the problem, but again, only in part. Some might see a few immediate issues. One is that not all site are sufficiently dynamic to know what content they actually contain. To an extent this might be addressable through extension to the webserver protocol such that a server would be aware, or become aware, of what content it contained. Another is that a site might in some instances be inclined to misrepresent what it contained. This may be hard for some to believe, but I'm given to understand it occasionally does occur. To help guard against this, there might be vetted indices, in which one or more third parties vouch for the validity of an index. These reputation-sources could of course themselves be assessed for accuracy. But if sites were responsible for reporting on what content they actually contained, and could be constrained to doing so accurately, a huge part of the overhead in creating independent search engine, and breaking the seach-engine monopoly, would be eliminated. One might imagine why certain existing gatekeepers over Web standards might oppose such an initiative. There would still remain other problems to solve within search space. It's possible to divide General Web Search into a set of specific problems: - Site crawling: this includes determining search targets, any exclusions from such lists, and performing the actual crawling. Self-indexing addresses part of this problem. - Indexing: Mapping of actual contents to keyword and query terms which might address that content. - Ranking: Assigning a preference / deprecation to specific sites. This is essentially a trust / reputation assessment, with a canonicity / authenticity assessment (e.g., where did a specific item or document first appear). - SEO: This is the Red Queen's Race issue in addressing insincere / malicous actors. Strong and durable penalties for abuse, and long-term reputational accrual, should be useful here. - Query interpretation: There's a considerable art to figuring out what a question actually means. In some cases queries should be taken strictly verbatim. Quite often, however, interpretation is necessary. How those alternatives are posed might vary, with an option not often employed presently being to suggest a range of potential interpretations or related queries which might produce better results for specific query scenarios. - Presentation: This is generation of the serch engine result page itself, incorporating several of the other considerations listed, but also addressing usability, accessibility, clarity, and other concerns. - Revalidation: As the editors of the Hitchiker's Guide observed, the Universe is not static, and circumstances change. Revalidating, revisiting, and revising results and reputational assessments is necessary. - Monetisation/Funding: I'm partial to a public goods model, or perhaps a farebox role via ISPs, pro-rated to general income/wealth within a region. Advertising, as a famous Stanford research paper prophetically observed, forces disallignment with searchers' interests and objectives.
- tinodb 4y agoHow should I pronounce this search engine? I know naming is hard, but if you want something to be easily adopted, having a sticky and pronounceable name is paramount!
- keynesyoudigit 4y agoI hope I'm not too late and this doesn't get buried - anyone interested should check out https://www.findhelp.org/ https://www.findhelp.org/ ! I work here and we are super hiring for engineers :) Edit - ah, he means the search engine should be a non-profit. Not what I thought he meant.