8 ms·
These solutions don't answer any of the fundamental problems with Google: - who pays for the service (ads? users pay? Average user will never use a paid servic
by boomer918 5y ago
These solutions don't answer any of the fundamental problems with Google:
- who pays for the service (ads? users pay? Average user will never use a paid service if a free one is available)
- how to resist attacks against the algorithm (Google has been fighting spam for decades)
- how to personalize without invading privacy, e.g. Google had an option to search through your email in Google search...it's gone now, I wonder why?
- chaostheory 5y agoAdding on to this, customization is nice but customization is not why DuckDuckGo isn't as good as Google. The reason nothing is as good as Google is because Google indexes way more content than every other service that I'm aware of
- prawn 5y agoIs that the key though? There's a lot of stuff I'd be quite happy if they didn't index!
- chaostheory 5y agoImo Yes, because that’s the main difference between Google search and everyone else. Out index Google, and it’s possible to beat. That said, I do like the feature of ignoring entire domains. Google used to allow that
- bobajeff 5y agoLately Google has gotten so bad that I've even occasionally brought up Bing (which I've often referred to as the Zune of search engines) and gotten better results. It's looking more and more to me like Google has stopped trying to improve things and are simply milking their dominance for all it's worth.
- _jal 5y agoAlso important for anyone actually thinking of taking Google on, very few of the features listed are things Google can't easily do, too. Attacking their strengths is crazy. You better have something both crazy good and hard to replicate by someone with more money than god. Whatever replaces Google will be doing something that Google can't without causing them other problems. The first thing that comes to mind is make them choose traffic vs. advertisers (I don't know, if I had an idea of how to, I would not be writing this), but they're big enough that other wedges could start chipping away at their margins.
- mkmk3 5y agoI think this comment is a bit strange in the present, considering search engines like duckduckgo, which is basically Bing promoted with a "we don't track" advertising campaign (also hashbangs are pretty cool). DDG is not at google numbers, I know, but you don't need google numbers to make money. I don't think privacy is a very special angle to advertise from either, promising to remove amazon-affiliate blog-spam from results for example, would be a major feature in this space as far as I'm concerned. Being able to edit searches, and potentially gain some intuition for how the search space is set up, might be a much more significant feature, depending on how people take to it. It might flop but atm I'm excited to check it out
- _jal 5y ago> but you don't need google numbers to make money The article is titled "The Next Google." I was responding to that, not "A Profitable Also-Ran".
- sanxiyn 5y agoWhile "The Next Google for Wall Street" is one interpretation of "The Next Google", I am more interested in "The Next Google for me".
- darinf 5y agoActually, you are spot on. One simple feature of the Neeva app is that it shows inline search results as you type into the URL bar. This is because we aren't trying to show you ads, so we don't need you to visit the search results page (where Google and others show you those ads). We just show you the results straight away in the suggest experience. Now, this isn't going to show you everything you care about and you can still click to see the search results page. It is just handy to be able to quickly get to where you are trying to go and especially if it is likely to match what you are looking for (e.g., a wikipedia link). This is something Google cannot bring itself to do because it would be cost way too much in terms of lost ads revenue. There are other examples like this where Google and other ad-supported search engines just can't innovate, can't change the search experience. The current way of searching is too lucrative and there is too much business inertia around it. That's why Neeva is interesting and why I left Google to join and help :)
- sanxiyn 5y agoKagi is "users pay". Yes, average users won't pay, but I don't see how that matters to me as a Kagi user.
- sillysaurusx 5y agoA paid search engine? Bold. I'd pay if it could do anything close to what the old google code search could do. I miss it every day.
- dx034 5y agoThe cost per search query is incredibly small. You could probably have <1% premium users. If your search engine is used by a billion people a day, 10 million paying users are still enough to pay a few thousand people developing your product (depending on your location). I believe products can work that way, offer premium service and features for the very few that need it and a basic service to anyone else. In the end, the free tier is cheaper than what you'd have to spend on marketing otherwise.
- josh_p 5y agoIt's free while they're in beta. There's a "waitlist" but put your email on it and you'll get an invite within a week. Give it a try. I've been using it for a couple weeks now on my work laptop and for programming-related searches its been great so far. And the usual annoyances that show up at the top of google and DDG don't show up on Kagi (geeks4geeks, etc).
- fallat 5y agoIt's actually so good I plan to pay when they start charging.
- marginalia_nu 5y agoSeems the first of these can be solved by reducing the scope. Do you really need a data center to run a search engine? Overall it seems very rare anyone ever considers this an engineering problem. Really, what's stopping you from running a search engine?
- sdoering 5y agoReally? Or are you the one I should have refrained from feeding. But if you must know: First you need to collect a lot of content from the internet. From many different sites. With very different types of code structure. Broken html. More often than not behind some SPA JS code. Behind robots.txt files and bot protection efforts. So the first problem to solve would be building a crawler at scale. That is able to crawl anything your users might want to visit but don't know of yet. Then storage and retrieval. You need to store and update all this content your crawler collected. You need to enrich it with meta data and organize it for efficient retrieval. So that you can surface it to your users when they use your search engine. Indexing, structure, build g connections between content pieces. A lot of interesting things to think about. Then there is the front end. Make it easy to search, to refine. Surface relevant content for search queries. OH maybe I forgot, but you probably need to do a bit of engineering to make your system understand the users' search intent. This is relatively straightforward for a limited search and document space up to a few million entries in your DB. A few million documents should be doable with off the shelf parts. Bigger than that. I would applaud you if done with orders of magnitude lower than Google. Anyone would.
- sanxiyn 5y agoSince you didn't seem to notice the username: you are replying to a person who developed a search engine alone. (So be prepared to applaud.)
- marginalia_nu 5y agoAll of this is a long series of solvable problems. I should know, I've dabbled in solving most of them. This is why I suggest actually taking a stab at it before you dismiss it as impossible. There are some problems that aren't as big as they seem. Parts of an SPA can't be reliably linked to anyway even if you find interesting text there, so you can just leave them out of the index. Likewise, there isn't as great of a need to keep a fresh index as it may seem. The odds of a document changing is proportional to how frequently it changes. This is a bit of a paradox, where even if you crawl really aggressively, the most frequently changing documents will still always be out of date. Most documents are relatively stable over time. You can actually use how often you see changes to a document or website to modulate how often you crawl it. The bad HTML is quite manageable. You really just need to flatten the document to get at the visible text. Even with really broken formatting, that's manageable. The storage demands are also not as bad as you might think (most documents are tiny, sub 10 Kb), there are ways to lessen the blow on top of that. Both text and indexes can compress extremely well. Since you're paying for disk access by the block, you might as well cram more stuff into a block. Most of the crawling concerns, in general, can be gotten around by starting off with Common Crawl (even if I do my own crawling, which also is finnicky but manageable). > This is relatively straightforward for a limited search and document space up to a few million entries in your DB. A few million documents should be doable with off the shelf parts. Right, so shouldn't the question be how to find the documents that are even candidates for being search results? Most documents are not ever going to be relevant to any query ever. Get rid of that noise and your hardware goes a lot longer. I'm running a search engine on consumer hardware out of my living room that can index 100 million documents. Go a bit higher budget than a consumer PC, and you've got 5 billion. That goes a long way.
- RcouF1uZ4gsC 5y ago> - who pays for the service (ads? users pay? Average user will never use a paid service if a free one is available) - how to resist attacks against the algorithm (Google has been fighting spam for decades) The solution for o both of these might actually be a paid service. If you have a paid service, there is a possibility of it being profitable with much fewer users. As an example, let’s say you have 1,000,000 users at $10/month, that is a $10,000,000/month which might be enough to run the service and provide a comfortable profit. With regards to the spam issue, the fact that you have a small user base would be to your advantage. Because there are so many Google users, it is in websites’ economic interests to spend money to try to game the algorithms. With much fewer users, your paid search users may not be worth it for the sites to spend money trying to game your algorithms.
- avsteele 5y agoIt will still pay to put ads in to paying customer's feeds. It's more or less inevitable if the customers tolerate it. If you think that's impossible I'd point to how streaming services are now serving up ever more adds.
- Nextgrid 5y agoThe spam issue can trivially be addressed by implementing actual penalties for rule-breakers. If it takes a long time to acquire a good reputation & ranking on the search engine, you're unlikely to risk it by doing something nasty in fear of your domain, keyword or brand name being banned for a long time.
- aaomidi 5y agoI do think there's actually some space opening up for paid services. From what I'm seeing, if you could create a bot free eco system, people will pay for it. The question is "can you make it bot free". This is gonna be the next trillion dollar company.
- Nextgrid 5y agoRaising the cost of spam would be a good first step. At the moment, spamming Google seems to be trivial with no long-term penalties if you get caught doing something nasty. A simple rule (manually enforced on a case-by-case basis) that would ban your brand/domain for a year if you get caught breaking the rules would get Pinterest into compliance from day 1 for example. Using ads/analytics/affiliate links as a negative ranking signal would make a lot of blogspam/listicles/clickbait disappear if their only funding method immediately makes them rank much lower below where they are no longer profitable.
- sshumaker 5y agoThis would be easily exploitable by a competitor. For example, search engines (used to) rank back links - that is other domains pointing to your domain. Some bad actors took advantage of this by creating rings of sites that voted each other up. Google responded by punishing the behavior. Then, competitors started taking advantage of this punishment by creating a network of sites that backlinked to a competitor, so they would get punished instead. This isn’t a hypothetical example - Google actually includes in their webmaster tools a “disavow links” capability so sites can avoid getting punished for bad actors trying to make them look bad. But you can imagine if the penalties were even more severe other folks may get caught up in an unforgiving dragnet with no judge or jury and no way to appeal. My main point is that people will find ways to game the system, and usually sharp edges (“harsh punishments”) on any system will be taken advantage of by actors, and unfairly penalize others.
- Nextgrid 5y agoAgreed, I'm not saying this is the end-game or that it will be perfect. But a simple rule (that's actually enforced) saying that you are forbidden to serve a different experience to the Google bot vs a normal visitor would take care of Pinterest for example, and they're not even doing that despite it being a major complaint especially in tech-circles where Googlers no doubt lurk.
- sjg007 5y agoI would pay for an app that searches my stuff and provides some kind of intelligent agent for web knowledge.
- Nextgrid 5y agoAverage users may not pay, but specialized users may pay and pay more than enough to subsidize some sort of free tier. Not to mention, if free search engines keep devolving into an endless sea of spam, people may have no choice but to start paying. There's plenty of things out there people pay for not necessarily by choice but because there's nothing else out there that would accomplish the task at hand.
- scim-knox-twox 5y ago> Average users may not pay, but specialized users may pay and pay more than enough to subsidize some sort of free tier. If I can, I'm happy to pay, so others doesn't need to. I don't understand why I'm in minority and most of the people thinking only about themselves. I'm happy to pay because for many years I also had to use free services, paid by others (or ads, but ads I'm blocking however I can).