4 ms·
One thing that consistently blows my mind is how readily an external search engine creates a better index, with more relevant recall, by scraping the published
by carterage 8y ago
One thing that consistently blows my mind is how readily an external search engine creates a better index, with more relevant recall, by scraping the published static pages of Wikipedia, such that it outperforms Wikipedia's own search,
Wikipedia should be able to search itself better than any external entity, but cannot. Wikipedia, in effect, is blind to its own data, and can deliver less insight into it's own content than multiple external Search entities.
If I use https://en.wikipedia.org/wiki/Special:Search https://en.wikipedia.org/wiki/Special:Search for anything, I get weaker results, than if I use search engines to locate wikipedia articles according to the same input string.
There's something wrong with that.
- barking 8y agoThe same is true for every forum I know of, in fact a lot of them now have a google search box that gives site specific results
- Larrikin 8y agoSearch isn't easy and I would fully expect a company that specializes in search to provide better search results than an entity that simply has search as a feature to their main product. Since Wikipedia provides all of their pages for indexing, unlike Facebook for example, a good search experience would be an MVP for any competent search company.
- carterage 8y agoThat's a good point. Put another way, were I to run a local copy of a wikipedia mirror, right here in my own bedroom, I'd be putting myself eye-to-eye with any wikipedia employee or volunteer. I'd have the same tools in my hands, as they do. I'd be one person, a PHP script hosted by an apache process, and the wikipedia data set. But, contrast that to Facebook, and the concept becomes more interesting. Facebook's search utilities are similarly unsatisfying, despite the massive resources of the company. This is probably not a blind spot, but more about the level of quality provided, for free-tier, unprivileged user utilities. Advertisers probably gain better reach, but perhaps without knowing exactly who they reach. Facebook's search tool can't be used to discern the capabilities or qualities of one or more of the indices Facebook uses to negotiate the landscape of their data.
- anonytrary 8y ago> There's something wrong with that. I disagree, I think Wikipedia is akin to a collective blog where people just dump information to. Indexing is hard enough that it becomes a full-time job to do it well, and that takes time and money. Search is so hard that external search is necessary for many websites to exist outside of a void. If we expect blogs to have their own searching functionality, we end up with a collection of disjoint webs, and that would basically kill the web as we know it.
- lloeki 8y agoThis seems typical of many built-in searches. It's halfway how things are indexed but also how things are presented: one of the worst offenders is ruby's documentation, where it's extremely cumbersome to get the latest version of a given item[0]. And that's a trivial search query. [0]: https://ruby-doc.com/search.html?q=Array https://ruby-doc.com/search.html?q=Array
- rileymat2 8y agoI don't see why this would be the case. Google has the advantage of seeing the content on pages with inbound links.