17 ms·
Wow, that's awesome. Great work! For a simple test, I searched "fall of the roman empire". In your search engine, I got wikipedia, followed by academic talks
by snakeboy 5y ago
Wow, that's awesome. Great work!
For a simple test, I searched "fall of the roman empire". In your search engine, I got wikipedia, followed by academic talks, chapters of books, and long-form blogs. All extremely useful resources.
When I search on google, I get wikipedia, followed by a listicle "8 Reasons Why Rome Fell", then the imdb page for a movie by the same name, and then two Amazon book links, which are totally useless.
- hdjjhhvvhga 5y agoAs long as few people use it, it will be great. Rest assured that the moment it becomes popular, the people who want to game it will appear.
- jandrese 5y agoThis sort of optimization is why simple recipes are typically found at the end of a rambling pointless blog post now. Still, the best way to break SEO is to have actual competition in the search space. As long as SEO remains focused on Google there is an opportunity for these companies to thrive by evading SEO braindamage.
- xtracto 5y agoThat's why I use Saffron [1], it magically converts those sites into a page in my recipe book. I found it when the developer commented here in HN. Also, a lot of cooking website have started to add a link with "jump to recipe" functionality allowing you to skip all the crap. [1] https://www.mysaffronapp.com/ https://www.mysaffronapp.com/
- Funes- 5y agoThere's also https://based.cooking https://based.cooking.
- eigengrau5150 5y agoRun by Luke Smith, an admitted neo-reactionary and possible white supremacist who writes like a 4chan reject.
- ggggtez 5y agoI've noticed this pattern start to pop up elsewhere. I've started to train my skimming skills, skipping a paragraph or two at a time to get past the fluff. Like an article about some current event will undoubtedly begin with "when I was traveling ten years ago...".
- WorldMaker 5y agoThat sort of recipe blog hasn't happened just for SEO. It's also a bit of a "two audiences" problem: if you are coming to that food blogger from a search you certainly would prefer the recipe first and then maybe any commentary on it below if the recipe looks good. If you are a regular reader of that food blogger you are probably invested in the stories up top and that parasocial connection and the recipes themselves are sometimes incidental to why you are a regular reader. You see some of that "two readers" divide sometimes even in classic cookbooks, where "celebrity" chefs of the day might spend much of a cookbook on a long rambling memoir. Admittedly such books were generally well indexed and had table of contents to jump right to the recipes or particular recipes, but the concept of "long personal ramble of what these recipes mean to me" is an old one in cookbooks too.
- run-types 5y agoI've basically never been taken to a recipe without a rambling preamble from Google. While food blogs may serve two audiences, a long introduction seems to be a requirement to appear in the top Google search results.
- WorldMaker 5y agoPersonally, I think that has a lot more to do with the fact that Google killed the Recipe Databases. There did used to be a few startups that tried to be Recipe Aggregators with advertising based business models, that would show recipes and then link to source blogs and/or cookbooks, and in the brief period where they existed Google scraped them entirely and showed entire recipes on search results and ate their ad revenue out from under them.
- tomrod 5y agoThat is a really bad thing by Google. Their core business is not recipes.
- kwertyoowiyop 5y ago
- SerLava 5y agoThat's not really for SEO, which favors readily accessible information. That's ads. When mobile users have to scroll past 10 add, theyll click on some of them and make the blog money.
- YeGoblynQueenne 5y ago>> This sort of optimization is why simple recipes are typically found at the end of a rambling pointless blog post now. I continue to be curious about this kind of complaint. If all you want is a recipe list, without any of the fluff, why would you click on a link to a blog, rather than on a link to a recipe aggregator? Foodie blogs exist specifically for the people who want a foodie discussion and not just an ingredients' list. Is it because blogs tend to have better recipes overall? In that case, isn't there a bit of entitlement involved in asking that the author self-sacrificingly provides only the information that you want, without taking care of their own needs and wants, also?
- Loughla 5y agoIt's the same thing that people always complain about. This thing is not in a format that I like, so it must be not what anyone likes. If you want JUST recipes, pay money instead of just randomly googling around. America's test kitchen has a billion, vetted, and really good recipes. That solves that problem.
- joegahona 5y agoI think the complaint is that those blogs rank higher than nuts-and-bolts recipes now. It wasn't that way a few years ago. Yes, scrolling down the results to Food Network or Martha Stewart or whatever is possible, as is going directly to those sites and using their site search, but it's noticeable and annoying.
- YeGoblynQueenne 5y agoNot my experience. For a very quick test, I searched DDG for "omelette recipe, "carbonara recipe" and "peking duck recipe" (just to spice it up a bit) and all my top results are aggregators. Even "avgolemeono recipe" (which I'd think is very specialised) is aggregators on top. To be honest, I don't follow recipes when I cook unless it's a dish I've never had before. At that point what I want is to understand the point of the dish. A list of ingredients and preparation instructions don't tell me what it's supposed to taste and smell like. The foodie blogs at least try to create a certain... feeling of place, I suppose, some kind of impression that guides you when you cook. I wouldn't say it always works but I appreciate the effort. My real complaint with recipe writers is that they know how to cook one or two dishes well and they crib the rest off each other so even with all the information they provide, you still can't reliably cook a good meal from a recipe unless you've had the dish before. But that's my personal opinion.
- zerd 5y agoIt's also because that's a way of trying to copyright protect recipes, which are normally not copyright protected. > “Mere listings of ingredients as in recipes, formulas, compounds, or prescriptions are not subject to copyright protection. However, when a recipe or formula is accompanied by substantial literary expression in the form of an explanation or directions, or when there is a combination of recipes, as in a cookbook, there may be a basis for copyright protection.”
- JohnFen 5y agoBut that copyright protection only extends to the literary expression. The recipe itself is still not covered by copyright, even if accompanied by an essay.
- Aeolun 5y agoSearching for ‘chocolate’ on this search engine turned up a surprisingly large amount of chocolate based recipes.
- Nextgrid 5y agoI don't think the existing media-heavy websites are gaming Google to rank higher. It's that Google itself prefers media heavy content; they don't have to "game" anything. I also think a search engine like this would be quite hard to game. An ML-based classifier trained on thousands of text-heavy and media-heavy screenshots should be quite robust and I think would be very hard to evade, so the "game" will become more about how identify the crawler so you can serve it a high-ranking page while serving crap to the real users, and it seems fairly easily to defeat if the search engine does a second pass using residential proxies and standard browser user agents to detect this behavior (it could also threaten huge penalties like the entire domain being banned for a month to even deter attempts at this).
- fragmede 5y agoWith the advances in text generation by machines that looks, but isn't quite accurate (aka GPT-3), seems like it would be easily gamed (given access to GPT-3). Even without GPT-3, if the content being prioritized is mere text, I'm sure that for a pile of money, I could generate something that looks like Wikipedia, in the sense that it's a giant pile of mostly text, but it would make zero sense to a human reader. (Building an SEO farm to boost ranking of not-wikpedia is left as an exercise for the reader.)
- the_other 5y agoIf there were a wider variety of popular search engines, with different ranking criteria, would sites begin to move away from gaming the system? Surely it would be too hard to game more than one search engine at a time?
- Nasrudith 5y agoIt would be a matter of numbers anyway about which they optimize for. A/B testing is already in place and doesn't care about where it comes from, just which one does better.
- new_guy 5y ago> the people who want to game it will appear. So just add human review to the mix, if a site is obviously trying to game the system (listicles, seo spam etc) just drop and ban them from the search index.
- hdjjhhvvhga 5y agoCongratulations, you've just invented negative SEO.
- phendrenad2 5y agoThere should be some perfect balance where this search engine is N% as popular as Google, where Google soaks up all of the gamifiers, but this search engine is still popular enough to derive revenue and do ML and other search-engine-useful stuff.
- marginalia_nu 5y agoSpecialization mostly a problem in monocultures. If you almost only plant wheat, you are going to end up with one hell of a pest problem. If you almost only have Windows XP, you are going to have one hell of a virus problem. If you almost only have SearchRank-style search engines (or just the one), you are going to have one hell of a content spam problem. Even though they have some pretty dodgy incentives, I don't think google suffers quality problems because they are evil, I think ultimately they suffer because they're so dominant. Whatever they do, the spammers adapt almost instantly. A diverse ecosystem on the other hand limits the viability of specialization by its very nature. If one actor is attacked, it shrinks and that reduces the opportunity for attacking it.
- adventured 5y agoI did a search for "George Washington" First result after Wikipedia: "Radiophone Transmitter on the U.S.S. George Washington (1920) In 1906, Reginald Fessenden contracted with General Electric to build the first alternator transmitter. G.E. continued to perfect alternator transmitter design, and at the time of this report, the Navy was operating one of G.E.'s 200 kilowatt alternators http://earlyradiohistory.us/1919wsh.htm http://earlyradiohistory.us/1919wsh.htm " Another result in the first few: " - VANDERBILT, GEORGE WASHINGTON PH: (800) ###-#233 FX: (#03) 641-5###. https://www.ScottWinslow.com/manufacturer/VANDERBILT_GEORGE_WASHINGTON/2182 https://www.ScottWinslow.com/manufacturer/VANDERBILT_GEORGE_... " And just below that terrible result: "I Looked and I Listened -- George Washington Hill extract (1954) Although the events described in this account are undated, they appear to have occurred in late 1928. I Looked and I Listened, Ben Gross, 1954, pages 104-105: Programs such as these called for the expenditure of larger sums than NBC had anticipated. It be http://earlyradiohistory.us/1954ayl2.htm http://earlyradiohistory.us/1954ayl2.htm " Dramatically worse than Google. --- Ok, how about a search for "Rome" then? Surely it'll pull some great text results for the city or the ancient empire. First result after Wikipedia: "Home | Rome Daily Sentinel Reliable Community News for Oneida, Madison and Lewis County http://romesentinel.com/ http://romesentinel.com/" The fourth result for searching "Rome": "Glenn's Pens - Stores of Note Glenn's Pens, web site about pens, inks, stores, companies - the pleasure of owning and using a pen of choice. Direcdtory of pen stores in Europe. http://www.marcuslink.com/pens/storesofnote/roma.html http://www.marcuslink.com/pens/storesofnote/roma.html" Again, dramatically worse than Google. --- Ok, how about if I search for "British"? First result after Wikipedia: "BRITISH MINING DATABASE British_Mining_Database http://www.users.globalnet.co.uk/~lizcolin/bmd.htm http://www.users.globalnet.co.uk/~lizcolin/bmd.htm " And after that: "British Virgin Islands Many of these photos were taken on board the Spirit of Massachusetts. The sailing trip was organized by Toto Tours. Images Copyright © Lowell Greenberg Home Up Spring Quail Gardens Forest Home Lake Hodges Cape Falcon Cape Lookout, Oregon Wahkeena http://www.earthrenewal.org/british_virgin_islands2.htm http://www.earthrenewal.org/british_virgin_islands2.htm" Again, far off the mark and dramatically worse than Google. I like the idea of Google having lots of search competition, this isn't there yet (and I wouldn't expect it to be). I don't think overhyping its results does it any favors.
- JasonFruit 5y ago
- klntsky 5y agoHowever, when searching for "haskell type inference algorithm" I get completely useless results.
- klntsky 5y agoSince it does not use synonyms, it looks like it is unable to answer "how's that thing called"-queries.
- burkaman 5y agoThat query is too long apparently. But if you shorten to "haskell type inference", I think it delivers on its promise: > If you are looking for fact, this is almost certainly the wrong tool. If you are looking for serendipity, you're on the right track. When was the last time you just stumbled onto something interesting, by the way?
- marginalia_nu 5y agoThe search engine doesn't do any type of re-ordering or synonym stuff, it only tires to construct different N-grams from the search query. So if you for example compare "SDL tutorial" with "SDL tutorials". On google you'd get the same stuff, this search engine, for better or worse doesn't. This is a design decision, for now anyway, mostly because I'm incredibly annoyed when algorithms are second-guessing me. On the other hand, it does mean you sometimes have to try different searches to get relevant results.
- leephillips 5y agoI like this design decision. It pays you back for choosing your search terms carefully.
- mananaysiempre 5y agoI’m not against a stemmer, actually, just against the aggressive concordances (?) that Google now employs, like when it shows me X in Banach spaces (the classical, textbook case) when I’m specifically searching for X in Fréchet spaces (the generalization I want to find but am not sure exists); of course Banach spaces and Fréchet spaces are almost exclusively encountered in the same context, but it doesn’t mean that one is a popular typo for the other! (The relative rarity of both of these in the corpus probably doesn’t help. The farcical case is BRST, or Becchi-Rouet-Stora-Tyutin, in physics, as it is literally a single key away from “best” and thus almost impossible to search for.) On the other hand, Google’s unawareness of (extensive and ubiquitous) Russian noun morphology is essentially what allowed Yandex to exist: both 2011 Yandex and 2021 Google are much more helpful for Russian than 2011 Google. I suspect (but have not checked) that the engine under discussion is utterly unusable for it. English (along with other Germanic and Romance languages to a lesser extent) is quite unusual in being meaningfully searchable without any understanding of morphology, globally speaking.
- hn_throwaway_99 5y agoI had the exact opposite experience. I searched the site for "java", got a Wikipedia link first (for the island, not the programming language), and the 2nd result was to a random JEP page, and all the rest of the results were random tidbits about Java (e.g. "XZ compression algorithm in Java). Didn't get any high level results pointing to an overview of the language, getting started guides, etc.
- withinboredom 5y agoYou need to use some old school search techniques and search for “Java overview”
- _wldu 5y agoI'm not sure that's a bad thing.
- rovr138 5y agowell, they're results to java related items... What kind of links where you expecting to find?
- eterevsky 5y agoImagine if you were looking for the movie.
- MisterTea 5y agoThe you'd use a different search engine. Why does everything have to be a Swiss Army knife?
- zozbot234 5y agoOr you could just search for 'rome movie'. Though for more complex disambiguation you would need to resort to, e.g. schema.org descriptions (which are supported by most search engines, and the foundation for most "smart" search result snippets).
- eterevsky 5y agoThat's a fair point. This engine would be useful if you need grep over internet (by without regexes), i.e. when you want to find the exact phrases. But that's a relatively narrow use case.
- neltnerb 5y agoImagine including the search term "movie".
- yreg 5y agoThat doesn't do anything useful.
- lucideer 5y agoI tend to prefer Wikipedia for movies. The exception is actor headshots if I'm trying to identify someone, which Wikipedia lacks for licensing reasons, but otherwise Wikipedia tends to be better than IMDB for most needs. Wikipedia has an IMDB link on every article anyway. Another need I guess might be reviews, for which RT or MC are better than IMDB: not sure if either of those two will fare better than IMDB in this search engine but again Wiki has links out (in addition to good reception summaries)
- foofoo4u 5y agoGood comparison. Reminds me of an analogy I like to make of today's web, which is it feels like browsing through a magazine store — full of top 10s, shallow wow-factoids, and baity material. I genuinely believe terrible results like this are making society dumber.
- wwweston 5y agoIt's also possible that it's the other way around: a certain "common denominator" + algorithms that chase broad engagement = mediocre results. The real trick would be some kind of engine that can aim just above where the user's at.
- rchaud 5y agoThe context matters. I'd happily read "Top 10" lists on a website if the site itself was dedicated to that one thing. "Top 10 Prog Rock albums", while a lazy, SEO-bait title, would at least be credible if it were on a music-oriented website. But no, these stories all come from cookie-cutter "new media" blog sites, written by an anonymous content writer who's repackaged Wikipedia/Discogs info into Buzzfeed-style copy writing designed to get people to "share to Twitter/FB". No passion, no expertise. Just eyeballs at any cost.
- foofoo4u 5y agoThis got me thinking that maybe one of the other big reasons for this is that the algorithms prioritize newer pages over older pages. This produces the problem where instead of covering a topic and refining it over time, the incentive is to repackage it over and over again. It reminds me of an annoyance I have with the Kindle store. If I wanted to find a book on, let's say, Psychology, there is no option to find all-time respected books of the past centenary. Amazon's algorithms constantly push to recommend the latest hot book of the year. But I don't want that. A year is not enough time to have society determine if the material withstands time. I want something that has stood the test of time and is recommended by reputable institutions.
- amenod 5y ago
- SPBS 5y agoCool, it appears that the trend towards JS may be causing self-selection -- if a page has a high amount of JS, it is highly unlikely to contain anything of value.
- dv_dt 5y agoIf one could create an metric of ad to content ratio from the js used, I would guess that would be a nice differentiator too.
- sjtindell 5y agoTrue. Unfortunately many large corporate websites through which you pay bills, order tickets, etc. are becoming infested with JS widgets and bulky, slow interfaces. These are hard to avoid.
- artificial 5y agoConversely no software to install. Browser as a platform. Don’t have to boot to Windows to pay your bills with activex for example
- TeMPOraL 5y agoThe platform isn't the problem. The problem is with the amount of code that does something other than letting you "pay bills, order tickets, etc.".
- foxfluff 5y agoThe mostly JS-less web was fine, fast, and reliable 20 years ago and I never had ActiveX. I hear stories about Flash and ActiveX but I literally never needed these to shop or pay bills online. Payments also didn't require scripts from a dozen domains and four redirects..
- artificial 5y agoYup, and taking payments online was awful but privacy was more of a thing. In South Korea ActiveX was required until recently. https://www.theregister.com/2020/12/10/south_korea_activex_certs_dead/ https://www.theregister.com/2020/12/10/south_korea_activex_c...
- deleted 5y ago[deleted]
- Nition 5y agoThe Wikipedia link at the top is always given. It would maybe be good to make it a little clearer that it's not one of the true results.
- Ajef 5y agoI think this is just because of terms you have searched. In my test-searches Wikipedia has not come up once in first position (i think the highest was 3rd in the list). Here's what I've tried with a few variations: golang generics proposal, machine learning transformer, covid hospitalization germany [edit] formatting
- Nition 5y agoI think maybe it's a special insert at the top, but only if a Wikipedia page is found that matches you search term? I'm not sure now though.
- titzer 5y agoSearch engines whose revenue is based on advertising will ultimately be tuned to steer you to the ad foodchain. All the incentives are aligned towards and all the metrics ultimately in service of, profit for advertisers. Not in the 99% of people who can convinced to consume something by ads? Welp, screw you.
- phendrenad2 5y agoSearch engines should be something you pay for. Surely search engine powerusers can afford to pay for such a service. If Google makes $1 per user per month or something, that's not too high a bar to get over.
- wizzwizz4 5y agoIn which case, consider paying for something like Infinity: https://infinitysearch.co/ https://infinitysearch.co/
- titzer 5y agoSearch engines should be like libraries. At least some tiny sliver of the billions we spend on education and research should go to, you know, actually organizing the world's information and making it universally available.
- illys 5y agoI see another issue here: companies like Google prioritize information to 1) keep their users and 2) maximize their profit. If you move data organization to another type of organization (non-profit, state, universities - private or public), then the question of data prioritization becomes highly political. What should be exposed? What should not? What to put first? ... It is already, but to a smaller extend since money-making companies have little interest in data meaning, and high interest in the commercial value of their users.
- varjag 5y agoThe theoretical cap for this, if you include every human being on planet Earth, is 7 billion/month. This translates into $84 billion annual revenue. Google's revenue last year was 146 billion, and it operates not anywhere near the theoretical maximum. Most of that revenue is advertisement.
- rasz 5y ago>followed by a listicle "8 Reasons Why Rome Fell" but arent you curious about the 7th reason? it will surprise you!
- deleted 5y ago[deleted]
- resynth1943 5y agoYeah, Google tends to send a lot of junk back.
- purplefruit 5y agoWow I used "personality test" and actually got useful articles about personality theory. I'll actually use this!
- Siira 5y agoI tried some queries for Harry Potter fanfictions, and the results were pretty much completely unrelated. There weren’t that many results, either.
- marginalia_nu 5y agoI'm curious what you searched for. https://search.marginalia.nu/search?query=harry+potter+fanfiction&profile=default https://search.marginalia.nu/search?query=harry+potter+fanfi... This seems to return a pretty decent number of sites relating to that (as well as some sites not relating to that). The search engine isn't always great at knowing what a page is about, unfortunately. This seemed to return mostly relevant results https://search.marginalia.nu/search?query=%22harry+potter%22+fanfiction&profile=default https://search.marginalia.nu/search?query=%22harry+potter%22...
- Siira 5y agoYes, shorter queries return more relevant results. I think this was the first query that came to my mind: https://search.marginalia.nu/search?query=Best+%22harry+potter%22+fanfictions&profile=default https://search.marginalia.nu/search?query=Best+%22harry+pott...
- marginalia_nu 5y agoYeah, that's just not a type of query my search engine is particularly good at. It's pretty dumb, and just tries to match as much of the webpage against the query as it can. This used to be how all search engines worked, but I guess people have been taught by google that they should ask questions now, instead of search for terms. I wonder how I can guide people to make more suitable queries. Maybe I should just make it look less like google.
- psadri 5y agoInteresting choice of search topic. Are you trying to make an additional point?
- acchow 5y agoIf this search engine ever takes off, the listicle writers will just start optimizing for it too, right?
- dotancohen 5y agoMission accomplished, then.
- acchow 5y agoIf the goal was to remove modern web design, ok sure mission accomplished. If your goal was to create a search engine that ignored listicles and other fluff and instead got you meatier results like "academic talks" and such, then no.
- dotancohen 5y agoWhen a measure becomes a target, it ceases to be a good measure. https://en.wikipedia.org/wiki/Goodhart%27s_law https://en.wikipedia.org/wiki/Goodhart%27s_law
- HPsquared 5y agoI think it's a case where systems diversity can be an advantage. Much like how most malware was historically written for Windows and could be avoided by using Linux, the low-quality search engine bait is created for Google and can be avoided by using a different style of search engine.
- 1vuio0pswjnm7 5y agoNo one mentioned the "bonus" audio in the page source: https://www.youtube.com/watch?v=7fCifJR6LAY https://www.youtube.com/watch?v=7fCifJR6LAY
- marginalia_nu 5y ago;-)