20 ms·
I have been thinking about the same problem since a few weeks. The real problem with search engines is the fact that so many websites have hacked SEO that there
by rahulchhabra07 7y ago
I have been thinking about the same problem since a few weeks.
The real problem with search engines is the fact that so many websites have hacked SEO that there is no meritocracy left. Results are not sorted based on relevance or quality but by SEO experts' efforts at making the search results favor themselves. I can possibly not find anything deep enough about any topic by searching on Google anymore. It's just surface-level knowledge that I get from competing websites who just want to make money off pageviews.
It kills my curiosity and intent with fake knowledge and bad experience. I need something better.
However, it will be interesting to figure the heuristics to deliver better quality search results today. When Google started, it had a breakthrough algorithm - to rank page results based on number of pages linking to it. Which is completely meritocratic as long as people don't game for higher rankings.
A new breakthrough heuristic today will look something totally different, just as meritocratic and possibly resistant to gaming.
- JohnJamesRambo 7y agoGoodhart‘s Law in action. I wonder how we make a new measure that buys us more time? https://en.wikipedia.org/wiki/Goodhart's_law https://en.wikipedia.org/wiki/Goodhart's_law
- vanderZwan 7y agoThe backtick in your link broke it: https://en.wikipedia.org/wiki/Goodhart's_law https://en.wikipedia.org/wiki/Goodhart's_law (Were you on mobile and using a smart keyboard?)
- JohnJamesRambo 7y agoI manually added the comma on a mobile smart keyboard. :) Didn't know that doesn't work haha.
- zozbot234 7y agoThat damn ‘Smart Quotes’ misfeature is still causing havoc even after 30 years.
- klingonopera 7y agoNitpick: It's actually an apostrophe " ' ", not a backtick/grave accent " ` " or comma " , " :D https://en.wikipedia.org/wiki/Apostrophe https://en.wikipedia.org/wiki/Apostrophe https://en.wikipedia.org/wiki/Grave_accent#Use_in_programming https://en.wikipedia.org/wiki/Grave_accent#Use_in_programmin... https://en.wikipedia.org/wiki/Comma https://en.wikipedia.org/wiki/Comma
- JohnJamesRambo 7y agoOh yes sorry I’m full of mistakes today. Of course, not a comma!
- MayeulC 7y ago> A new breakthrough heuristic today will look something totally different, just as meritocratic and possibly resistant to gaming. I wonder how much of this could be obtained back by penalizing: 1. The number of javascript dependencies 2. The number of ads on the page, or the depth of the ad network This might start a virtuous circle, but in the end, this is just a game of cat-and-mouse, and website might optimize for this as well. What we might need to break this is a variety of search engines that uses different criteria to rank pages. I suspect it would be pretty hard, if not impossible, to optimize for all of them. And in any case, frequently change the ranking algorithms to combat over-optimization by the websites (as that's classically done against ossification for protocols, or any overfitting to outside forces in a competitive system).
- derefr 7y agoYou could even have all this under one roof: one common search spider that feeds this ensemble of different ranking algorithms to produce a set of indices, and then a search engine front end that round-robins queries out between the different indices. (Don’t like your query? Spin the algorithm wheel! “I’m Feeling Lucky” indeed.)
- zozbot234 7y agoThe Common Crawl is a thing already. Unfortunately, a "full" text crawl of the internets is a YUUUGE amount of data to manage, and I can't think of anything that could change that in the foreseeable future. That's why I think providing a federated Web directory standard, ala ODP/DMOZ except not limited to a single source, would be a far more impactful development.
- reaperducer 7y agoUnfortunately, a "full" text crawl of the internets is a YUUUGE amount of data to manage Maybe instead of a problem, there is an opportunity here. Back before Google ate the intarwebs, there used to be niche search engines. Perhaps that is an idea whose time has come again. For example, if I want information from a government source, I use a search engine that specializes in crawling only government web sites. If I want information about Berlin, I use a search engine that only crawls web sites with information about Berlin, or that are located in Berlin. If I want information about health, I use a search engine that only crawls medical web sites. Each topic is still a wealth of information, but siloed enough that the amount of data could be manageable to a small or medium-sized company. And the market would keep the niches from getting so small that they become useful. A search engine dedicated to Hello Kitty lanyards isn't going to monetize.
- ethbro 7y ago> However, it will be interesting to figure the heuristics to deliver better quality search results today. If only there were some kind of analog for effective ways to locate information. Like if everything were written on paper, bound into collections, and then tossed into a large holding room. I guess it's past the Internet's event horizon now, but crawler-primary searching wasn't the only evolutionary path to search. Prior to Google (technically: AdWords revenue funding Google) seizing the market, human-currated directories were dominant [1, Virtual Library, 1991] [2, Yahoo Directory, 1994] [3, DMOZ, 1998]. Their weakness was always cost of maintenance (link rot), scaling with exponential web growth, and initial indexing. Their strength was deep domain expertise. Google's initial success was fusing crawling (discovery) with PageRank (ranking), where the latter served as an automated "close enough" approximation of human directory building. Unfortunately, in the decades since we seem to have forgotten how useful hand-currated directories were, in our haste to build more sophisticated algorithms. Add to that that the very structure of the web has changed. When PageRank first debuted, people were still manually tagging links to their friends' / other useful sites on their own. Does that sound like the link structure we have in the web now? Small surprise results are getting worse and worse. IMHO, we'd get a lot of traction out of creating a symbiotic ecosystem whereby crawlers cooperate with human currators, both of whose enriched output is then fed through machine learning algorithms. Aka a move back to supervised web search learning, vs the currently dominant unsupervised. [1] https://en.m.wikipedia.org/wiki/World_Wide_Web_Virtual_Library https://en.m.wikipedia.org/wiki/World_Wide_Web_Virtual_Libra... , http://vlib.org/ http://vlib.org/ [2] https://en.m.wikipedia.org/wiki/Yahoo!_Directory https://en.m.wikipedia.org/wiki/Yahoo!_Directory [3] https://en.m.wikipedia.org/wiki/DMOZ https://en.m.wikipedia.org/wiki/DMOZ , https://www.dmoz-odp.org/ https://www.dmoz-odp.org/
- CM30 7y agoMixing human curation with crawlers is probably something that'd help with search results quality, but the issue comes in trying to get it to scale properly. Directories like the Open Directory Project/DMOZ and Yahoo's directory had a reputation for being slow to update, which left them miles behind Google and its ilk when it came to indexing new sites and information. This is problematic when entire categories of sites were basically left out of the running, since the directory had no way to categorise them. I had that problem with a site about a video game system the directory hadn't added yet, and I suspect others would have it for say, a site about a newer TV show/film or a new JavaScript framework. You've also got the increase in resources needed (you need tons of staff for effective curation), and the issues with potential corruption to deal with (another thing which significantly effected the ODP's usefulness in its later years).
- endothrowho333 7y agoThis is very much misguided. Many websites do have "hacked" (blackhat/shady) SEO, but these websites do not last long, and are entirely wiped out (see: de-ranked) every major algorithm update. The major players you see on the top rankings today do utilize some blackhat SEO, but it's not at a level that significantly impacts their rankings. Blackhat SEO is inherently dangerous, because Google's algorithm will penalize you at best when it finds out -- and it always does -- and at worst completely unlist your domain from search results, giving it a scarlet letter until it cools off. However, the bulk of all major websites primary utilize whitehat SEO, i.e "non-hacked," i.e "Google-approved" SEO to maintain their rankings. They have to, else their entire brand and business would collapse, either from being out-ranked or by being blacklisted for shady practices. Additionally, Google's algorithim hasn't changed much at all from pagerank, in the grand scheme of things. If you can read between their lines, the biggest SEO factor is: how many backlinks from reputable domains do you have pointing at your website? Everything else, including blackhat SEO, are small optimizations for breaking ties. Sort of like PED usage in competitive sports; when you're at the elite level, every little bit extra can make a difference. Google's algorithm works for its intended purposes, which is to serve pages that will benefit the highest amount of people searching for a specific term. If you are more than 1 SD from the "norm" searching for a specific term, it will likely not return a page that suits you best. Google's search engine based on virality and pre-approval. "Is this page ranked highly by other highly ranked pages, and does this page serve the most amount of people?" It is not based on accuracy, or informational-integrity -- as many would believe by the latest Medic update -- but simply "does this conform to normal human biases the most?" If you have a problem with Google's results, then you need to point the finger at yourself or at Google. SEO experts, website operators, etc. are all playing a game that's set on Google's terms. They would not serve such shit content if Google did not: allow it, encourage it, and greatly reward it. Google will never change the algorithm to suit outliers, the return profile is too poor. So, the next person to point a finger at is you: the user. Let me reiterate, Google's search engine is not designed for you; it is designed for the masses. So there is no logical reason for you to continue using it the way you do. If you wish to find "deep enough" sources, that task is on you, because it cannot be readily or easily monetized; thus, the task will not be fulfilled for free by any business. So, you must look at where "deep enough" sources lay: books, journals, and experts. Books are available from libraries, and a large assortment of them are cataloged online for free at Library Genesis. For any topic you can think of, there is likely to be a book that goes into excruciating detail that satisfies your thirst for "deep enough." Journals, similarly. Library Genesis or any other online publisher, e.g NIH, will do. Experts are even better. You can pick their brains and get even more leads to go down. Simply, find an author on the subject -- Google makes this very easy -- and contact them. I'm out of steam, but I really felt the need to debunk this myth that Google is a poor, abused victim, and not an uncaring tyrant that approves of the status quo.
- iovrthoughtthis 7y ago“When a measure becomes a target, it ceases to be a good measure” - Charles Goodhart
- graycat 7y agoGood to hear your concerns. > The real problem with search engines is the fact that so many websites have hacked SEO that there is no meritocracy left. I intend to announce the alpha test of my search engine here on HN. My search engine is immune to all SEO efforts. > I can possibly not find anything deep enough about any topic by searching on Google anymore. In simple terms my search engine gives users content with the meaning they want and in particular stands to be very good, by far the best, at delivering content with "deep" meaning. > I need something better. Coming up. > However, it will be interesting to figure the heuristics to deliver better quality search results today. Uh, sorry, it's not fair to say that my search engine is based on "heuristics". I'm betting on my search engine being successful and would have no confidence in heuristics. Instead of heuristics I took some new approaches: (1) I get some crucial, powerful new data. (2) I manipulate the data to get the desired results, i.e., the meaning. (3) The search engine likely has by far the best protections of user privacy. E.g., search results are the same for any two users doing the same query at essentially the same time and, thus, in particular, independent of any user history. (4) The search engine is fully intended to be safe for work, families, and children. For those data manipulations, I regarded the challenge as a math problem and took a math approach complete with theorems and proofs. The theorems and proofs are from some advanced, not widely known, pure math with some original applied math I derived. Basically the manipulations are as specified in math theorems with proofs. > A new breakthrough heuristic today will look something totally different, just as meritocratic and possibly resistant to gaming. My search engine is "something totally different". My search engine is my startup. I'm a sole, solo founder and have done all the work. In particular I designed and wrote the code: It's 100,000 lines of typing using Microsoft's .NET. The typing was without an IDE (integrated development environment) and, instead, was just into my favorite general purpose text editor KEdit. It's my first Web site: I got a good start on Microsoft's .NET and ASP.NET (for the Web pages) from Jim Buyens, Web Database Development, Step by Step, .NET Edition, ISBN 0-7356-1637-X, Microsoft Press. The code seems to run as intended. The code is not supposed to be just a "minimum viable product" but is intended for first production to peak usage of about 100 users a second; after that I'll have to do some extensions for more capacity. I wrote no prototype code. The code needs no refactoring and has no technical debt. While users won't be aware of anything mathematical, I regard the effort as a math project. The crucial part is the core math that lets me give the results. I believe that that math will be difficult to duplicate or equal. After the math and the code for the math, the rest has been routine. Ah, venture capital and YC were not interested in it! So I'm like the story "The Little Red Hen" that found a grain of wheat, couldn't get any help, then alone grew that grain into a successful bakery. But I'm able to fund the work just from my checkbook. The project does seem to respond to your concerns. I hope you and others like it. How should I announce the alpha test here at HN?
- BlackLotus89 7y ago> I have been thinking about the same problem since a few weeks. The real problem with search engines is the fact that so many websites have hacked SEO that there is no meritocracy left. That's not actually the problem described here. His problem is actually a bit deeper rootet since he specified the exact parameters of what he wants to see, but got terrible results. He specified a search for "site:reddit.com" but the resilts he got were ireelevant and worse than the results that he would have got when searching reddit directly. I don't say that SEO, sites that copy content and only want to genrate clicks and large sites that culminate everything are bad fkr the internet of today, but the level of results we get off of search engines today is with one word abysmal.
- mattmcknight 7y agoWrong. The site query worked. The issue is that there is no clear way to determine information date, as pages themselves change. Since more recent results are favored, SEO strategy of freshness throws off date queries. https://www.searchenginejournal.com/google-algorithm-history/freshness-update/ https://www.searchenginejournal.com/google-algorithm-history...
- BlackLotus89 7y agoWrong. In the article he elaborates. > As you can see, I didn’t even bother clicking on all of them now, but I can tell you the first result is deeply irrelevant and the second one leads to a 4 year old thread. He also wrote > At this point I visibly checked on reddit if there’ve been posts about buying a phone from the last month and there are. Duckduckgo even recognized the date to be 4 years old and reddit doesn't hide the age of posts. There are newer more fitting posts, but they aren't shown. And again a quote > Why are the results reported as recent when they are from years ago, I don’t know - those are archived post so no changes have been made. So your argument (also it really is a problem) in this case is a red herring. The problem lies deeper since google seems to be unable to do something as simple as extracting the right date and ddg ignores it. Also since all the results are years old it adds to the confusion why the results don't match the query. (He also wrote that the better matches were indeed indexed, but not found)
- AznHisoka 7y ago“I can possibly not find anything deep enough about any topic by searching on Google anymore. It's just surface-level knowledge that I get from competing websites who just want to make money off pageviews.” Is it possible that there is no site providing non fluffy content on your query? For a lot of niche subjects, there really are very few if any substantial content on that topic.
- Swizec 7y ago> Is it possible that there is no site providing non fluffy content on your query? For a lot of niche subjects, there really are very few if any substantial content on that topic. “Very few if any substantial” The problem is that google won’t even show me the very few anymore. It’s just fluff upon fluff and depth (or real insight at least) is buried in twitter threads and reddit/hn comments, and github issue discussion. I fear the seo problem has not only killed knowledge propagation, but also thoroughly disincentivized smart people from even trying. And that makes me sad.
- CriticalCathed 7y agoI mirror your sentiment. It used to be that you could use your google fu and you'd be able to find a dozen relevant forum posts or mail chains in plain text. It's much, much, harder to get the same standard of results. "Pasting stack traces and error messages. Needle in a haystack phrases from an article or book. None of it works anymore." Yeah, if I know where I'm looking (the sites) then google is useful since I can narrow it to that domain. But if I don't know where to look then I'm SOL. The serendipity of good results on Google.com is no longer there. And given the talent at google you have to wonder why.
- Swizec 7y agoThe devil is in this detail: Regular users don’t want those “weird looking” results. Normies prefer the fluff. And guess what: most users are normal. Us here on HN are weird:
- 7y ago
- erikbye 7y ago> It's just surface-level knowledge that I get from competing websites who just want to make money off pageviews. Can you give some examples of queries/topics? Not that I disagree, I often have the same problem, but have found ways to mitigate.
- edoceo 7y agoCan you elaborate some of these ways? I just have a big minus list of sites.
- erikbye 7y agoI would need to hear examples, queries or topics that results in solely superficial information, like OP stated.
- stevenicr 7y agoI too am asking people to start making lists of lame query returns. I have taken screen shots of some, even made a video about one... but a solid list in a spreadsheet perhaps would be helpful.. of course with results varying for different people / locations and month to month / year to year, having some screen shots would be helpful too. Not sure if there is a good program for snapping some screen shots and pulling some key phrases and putting it all together well...
- jacquesm 7y agoIt is only solvable for a short period of time. Then, when whatever replaces the current search is successful enough there will be an incentive to game the new system. So the only way to really solve this is by radical fragmentation of the search market or by randomizing algorithms.
- deleted 7y ago[deleted]
- zackees 7y agoThe real reason why search is so bad is that Google is downranking the internet. I should know - I blew the whistle on the whole censorship regime and walked 950 pages to the DOJ and media outlets. --> zachvorhies.com <-- What did I disclose? That Google was using a project called "Machine Learning Faireness" to rerank the entire internet. Part of this beast has to do with a secret Page Rank score that Google's army of workers assign to many of the web pages on the internet. If wikipedia contains cherry picked slander against a person, topic or website then the raters are instructed to provide a low page rank score. This isn't some conspiracy but something openly admitted by Google itself: https://static.googleusercontent.com/media/guidelines.raterhub.com/en//searchqualityevaluatorguidelines.pdf https://static.googleusercontent.com/media/guidelines.raterh... See section 3.2 for the "Expertise, Authoritativeness and Trustworthiness" score. Despite the fact that I've had around 50 interviews and countless articles written about my disclosure, my website zachvorhies.com doesn't show up on Google's search index, even when using the exact url as a query! Yet bing and duckduckgo return my URL just fine. Don't listen to the people who say that's its some emergent behavior from bad SEO. This deliberate sabotage of Google's own search engine in order to achieve the political agenda of the controllers. The stock holders of Google should band together in a class action lawsuit and sue the C-Level executives of negligence. If you want your internet search to be better then stop using Google search. Other search engines don't have this problem: I'm looking at qwant, swisscows, duckduckgo, bing and others. ~Z~
- erklik 7y ago> Despite the fact that I've had around 50 interviews and countless articles written about my disclosure, my website zachvorhies.com doesn't show up on Google's search index, even when using the exact url as a query! I am not sure what is happening but I directly searched for your website : zachvorhies.com on Google (in Australia, if that matters), which returned the website as the first result: https://i.imgur.com/Z7RTsuE.png https://i.imgur.com/Z7RTsuE.png
- DenisM 7y agoI'm in the US and Google does not display zachvorhies.com when searching for "zachvorhies.com".
- kccqzy 7y agoNot to totally detract from your point, but my previous experience with SEO people showed that some SEO strategies actually not only improve page ranking, but also actual usability. The first, and the most important perhaps was page load speed. We adopted a slightly more complicated pipeline on the server side, reduced the amount of JS required by the page, and made page loading faster. That improved both the ranking and actual usability. The second was that SEO people told us our homepage contained too many graphics and too few text, so search engines didn't quite extract as much content from our pages. We responded by adding more text in addition to the fancy eye-catching graphics. That improved both the ranking and actual accessibility of the site.
- stevenicr 7y agoI have noticed most HN comments with SEO in them take it as being bad bad bad and the reason for the death of good search, the need for powerful whatever.. I really wish everyone would qualify, and not just black-hat seo / whitehat - there are many types of SEO, often with different intentions. I understand there has been a lot of google koolaid (and others) about how seo is evil it's poisoning the web, etc... But now, or has it been a couple years how? google had a video come out saying an SEO firm is okay if they tell you it takes 6 months... they have upgraded their pagespeed tool which helps with seo, and were quite public about how they wanted ssl/https on everything and that that would help with google seo.. so there are different levels of SEO, someone mentioned an seo plugin I was using on a site as being a negative indicator, and I chuckled - the only thing that plugin does is try to fix some of the inherent obvious screwups of wordpress out of the box... things like no meta description which google flags on webmaster tools as multiple same meta descriptions.. also tries to alleviate duplicate content penalties by no-indexing archives or and categories or whatever. So there is SEO that is trying to work with google, and then there is SEO where someone goes out and puts comments on 10,000 web sites only for the reason of ranking higher.. to me that is kind of grey hat if it was a few comments, but shady if it's hundreds and especially if it's automated.. but real blackhat stuff - hacking other web sites and adding links.. or having a site that is selling breast enlarging pills and trying to get people who type in a keyword for britney spears.. that is trying to fool people. I have built sites with good info and had to do things to make them prettier for the ranking bot, but they are giving the surfer what they are looking for when they type 'whatever phrase'... I have also made web sites better when trying to also get them to show up in top results. So it's not always seo=bad, sometimes seo=good for the algorythm and the users. and sometimes it's madness - like extra fluff to keep a visitor on page longer to keep google happy like recipes - haha - many different flavors of it - and different intentions.
- creato 7y ago> I can possibly not find anything deep enough about any topic by searching on Google anymore. > It kills my curiosity and intent with fake knowledge and bad experience. I need something better. It's hard for me to take this seriously when wikipedia exists, and almost always ranks very highly in search results for searches for "knowledge topics". Between wikipedia and sources cited on wikipedia, I find the depth of almost everything worth learning about to be far greater than I can remember in, say, the early 2000s, which is seems like the "peak" of google before SEO became so influential. In general, I think there are a lot of people wearing rose tinted glasses looking at the "old internet" in this thread. The only thing that has maybe gotten worse is "commercial" queries like those for products and services. Everything else is leaps and bounds better.
- deleted 7y ago[deleted]
- dredmorbius 7y agoThe fact that Wikipedia exists, is frequently (though not always) quite good, has citations and references, and ranks highly or is used directly for "instant answers" ... ... still does nothing to answer the point that Web search itself is unambiguously and consistently poorer now than it was 5-10 years ago. Yes, I find myself relying far more on specific domain searches, either in the Internet sense, or by searching for books / articles on topics rather than Web pages. Because so much more traditionally-published information is online, this actually means the net of online-based search has improved, but not for the most part because of improved Web-oriented resources (Webpages, discussions, etc.), but because old-school media is now Web-accessible.
- sifar 7y agoThis. More and more, I have been finding that good books provide better learning than the internet. You search for quality books online, mostly through discussion forums as search fails here, or through following references of books and articles. Then spend time digesting them.
- joe_the_user 7y ago
- drpixie 7y agoFrom my end, it looks like google search is very strongly prioritising paid clients and excluding references to everything else. Try a search or view from maps, it shows a world that only includes google ad purchasers. Google has become a not very useful search - certainly not the first place I go when looking for anything except purchases. They've broken their "core business".
- mariushn 7y agoIt also favors itself. I searched for "work music". First 9 results are from youtube.
- eyegor 7y agoHere's an xkcd [0] inspired idea. We have several search engines, each with some level of bias. We're not looking to crawl the whole internet because we can't compete with their crawlers. However, we could make a crawler to crawl their results, and re-rank the top N from each engine according to our own metric. Maybe even expose the parameters in an "advanced" search. I'm assuming this would violate some sort of eula though. Any idea if someone has tried this approach? Edit: thinking more about this post's specific issue, I'm not sure what to do if all the crawlers fail. Could always hook into the search apis for github, reddit, so, wiki, etc. Full shotgun approach. [0] https://xkcd.com/927/ https://xkcd.com/927/
- Eridrus 7y agoIsn't this basically what DuckDuckGo does?
- Swalden123 7y agoI’ve found more and more I have reverted to finding books instead of searching to find deeper knowledge. The only issue is it is easy to publish low quality books now. Depending on the topic you are looking into often if a book stands the test of time it is a worthwhile read. With tech books you have to focus on the authors credentials.
- deleted 7y ago[deleted]
- kordlessagain 7y agoAn AI crawler is needed.
- arminiusreturns 7y agoI've often thought one approach, though one I wouldn't necessarily want to be the standard paradigm, would be exclusivity based on usefulness. So for example duckduckgo is still trying to use various providers to emulate essentially "early google without the modern google privacy violations", but when I start to think about many of the most successful company-netizens, one thing that stands out is early day exclusivity has a major appeal. So I imagine a search engine that is only crawling the most useful websites and blogs and works on a whitelist basis. Instead of trying to order search results to push bad results down, just don't include them at all or give them a chance to taint the results. It would have more overhead, and would take a certain amount of time to make sure it was catching non-major websites that are still full of good info ... but once that was done it would probably be one of the best search panes in existence. I have also thought to streamline this, and I know it's cliche, but surely there could be some ml analysis applied to figuring out which sites are SEO gaming or click-baiting regurgitators and weed them out. Just something I've been mulling over for a while now.
- Mauricio_ 7y agoAnd how do you determine which websites are good other than checking if they are doing seo? Is reddit.com good or bad? If a good site that does seo should it be taken out? And what if what you're searching exists only in a non good website? Isn't it better to show a result from a non good website than showing nothing?
- naravara 7y agoSomeone in a previous thread that I, unfortunately can’t remember, suggested that it might not just be the SEO but the internet that changed. Google used to be really good at ascertaining meaning from a panoply of random sources, but those sites are all gone now. The Wild West of random blogs and independent websites are basically dead in favor of content farms and larger scale media companies.
- ChuckMcM 7y agoIn the specific case of date based searches they are pretty difficult because of how pages are ranked. For a long time (and still to a large extent) Google ranks 'newer' pages higher than 'relevant' pages. At Blekko[1] there was a lot of code that tried to figure out that actual date of the document (be it a forum post, news article, or blog post). That date would often be months or years earlier than the 'last change' information would have you think. Sometimes its pretty innocuous, a CMS system updates every page with an updated copyright notice at the start of each year. Other times its less innocuous where the page simply updates the "related links" or side bar material and refreshes the content. It is still an unsolved ranking relevance problem where a student written, 3 month old description of how AM modulation works ranks higher than a 12 year old, professor written description. There isn't a ranking signal for 'author authority'. I believe it is possible to build such a system but doing so doesn't align well with the advertising goals of a search engine these days. [1] disclaimer I worked at Blekko.
- siruncledrew 7y agoPerhaps it’s not always a new heuristic that is needed, but a better way to manage the externalities around current/preceding heuristics. From a “knowledge-searching” perspective, at a very rudimentary level, it makes sense to look to sites/pages that are often cited (linked to) by others as better sources to rank higher up in the search results. It’s a similar concept to looking at how often academic papers are cited to judge how “prominent” of a resource they are on a particular topic. However, as with academia, even though this system could work pretty well for a long time at its best (science has come a long way over hundreds of years of publications), that doesn’t mean it’s perfect. There’s interference that could be done to skew results in one’s favor, there’s funneling money into pseudoscience to turn into citable sources, there’s leveraging connections and credibility for individual gain, - the list goes on. The heuristic itself in not innately the problem. The incentive system that exists for people to use the heuristic in their favor creates the issue. Because even if a new heuristic emerges, as long as the incentive system exists, people will just alter course to try to be a forerunner in the “new” system to grab as big a slice of the pie while they can. That’s a tough nut for google (or anyone) to crack. As a company, how could they actually curate, maintain, and evaluate the entire internet on a personal level while pursuing profitability? That seems near impossible. Wikipedia does a pretty damn good job at managing their knowledge base as a nonprofit, but even then they are always capped by amount of donations. It’s hard to keep the “shit-level” low on search results when pretty much anyone, anywhere, anytime could be adding in more information to the corpus and influencing the algorithms to alter the outcomes. It gets to a point where getting what you need is like finding a needle in a haystack that the farmer got paid for putting there.