5 ms·
Number of docs isn’t the limiting factor. I just searched for “stackoverflow” and the first result was this: https://www.perl.com/tags/stackoverflow/ https://w
by orf 9mo ago
Number of docs isn’t the limiting factor.
I just searched for “stackoverflow” and the first result was this: https://www.perl.com/tags/stackoverflow/ https://www.perl.com/tags/stackoverflow/
The actual Stackoverflow site was ranked way down, below some weird twitter accounts.
- saltysalt 9mo agoI don't weight home pages in any way yet to bump them up, it's just raw search on keyword relevance.
- orf 9mo agoSure, but the point is results are not relevant at all? It’s cool though, and really fast
- saltysalt 9mo agoI'll work on that adjustment, it's fair feedback thanks!
- direwolf20 9mo agoUnfortunately this is the bulk of search engine work. Recursive scraping is easy in comparison, even with CAPTCHA bypassing. You either limit the index to only highly relevant sites (as Marginalia does) or you must work very hard to separate the spam from the ham. And spam in one search may be ham in another.
- saltysalt 9mo agoI limit it to highly relevant curated seed sites, and don't allow public submissions. I'd rather have a small high-quality index. You are absolutely right, it is the hardest part!
- globular-toast 9mo agoWhat do you mean they're not relevant? The top result you linked contained the word stackoverflow didn't it? It's showing you exactly what you searched for. Why would you need a search engine at all if you already know the name of the thing? Just type stackoverflow.com into your address bar. I feel like Google-style "search" has made people really dumb and unable to help themselves.
- orf 9mo agothe query is just to highlight that relevance is a complex topic. few people would consider "perl blog posts from 2016 that have the stack overflow tag" as the most relevant result for that query.
- dredmorbius 9mo agoGoogle's entire (initial) claim-to-fame was "PageRank", referring both to the ranking of pages and co-founder Larry Page, which strongly prioritised a relevance attribute over raw keyword findings (which then-popular alternatives such as Alta Vista, Yahoo, AskJeeves, Lycos, Infoseek, HotBot, etc., relied on, or the rather more notorious paid-rankings schemes in which SERP order was effectively sold). When it was first introduced, Google Web Search was absolutely worlds ahead of any competition. I remember this well having used them previously and adopted Google quite early (1998/99). Even with PageRank result prioritisation is highly subject to gaming. Raw keyword search is far more so (keyword stuffing and other shenanigans), moreso as any given search engine begins to become popular and catch the attention of publishers. Google now applies other additional ordering factors as well. And of course has come to dominate SERP results with paid, advertised, listings, which are all but impossible to discern from "organic" search results. (I've not used Google Web Search as my primary tool for well over a decade, and probably only run a few searches per month. DDG is my primary, though I'll look at a few others including Kagi and Marginalia, though those rarely.) <https://en.wikipedia.org/wiki/PageRank https://en.wikipedia.org/wiki/PageRank> "The anatomy of a large-scale hypertextual Web search engine" (1998) <http://infolab.stanford.edu/pub/papers/google.pdf http://infolab.stanford.edu/pub/papers/google.pdf> (PDF) Early (1990s) search engines: <https://en.wikipedia.org/wiki/Search_engine#1990s:_Birth_of_search_engines https://en.wikipedia.org/wiki/Search_engine#1990s:_Birth_of_...>.
- saltysalt 9mo agoPageRank was an innovative idea in the early days of the Internet when trust was high, but yes it's absolutely gamed now and I would be surprised if Google still relies on it. Fair play to them though, it enabled them to build a massive business.
- marginalia_nu 9mo agoAnchor text information is arguably a better source for relevance ranking in my experience. I publish exports of the ones Marginalia is aware of[1] if you want to play with integrating them. [1] https://downloads.marginalia.nu/exports/ https://downloads.marginalia.nu/exports/ grab 'atags-25-04-20.parquet'
- pjc50 9mo agoConfluence search does this, for our intranet. As a result it's barely usable. Indexing is a nice compact CS problem; not completely simple for huge datasets like the entire internet, but well-formed. Ranking is the thing that makes a search engine valuable. Especially when faced with people trying to game it with SEO.