Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
deusu
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
deusu
10y ago
You can do your own search-index with 2.3 billion pages for about €300/month: https://deusu.org
2.
▲
Why DeuSu may succeed where Wikia Search failed
(deusu.org)
5 points
by
deusu
10y ago
|
0 comments
3.
▲
by
deusu
10y ago
It's alive and well. The TIOBE index still lists it ahead of Ruby, Swift, Objective-C, GoLang... And I started this software 20 years ago. Granted, a LOT of the software has changed since then. But I don't see a reason to throw aw
4.
▲
by
deusu
10y ago
Thank you!
5.
▲
by
deusu
10y ago
I don't know. But I do know that the end-of-year statistics from search-engines about what people searched for, are complete BS. I have such a list for the German DeuSu page: https://deusu.de/blog/2015-12-03-alle
6.
▲
by
deusu
10y ago
Yes. I have downloaded several data dumps, but haven't gotten around to import them yet.
7.
▲
by
deusu
10y ago
Currently €300/month. More details on https://deusu.org/donate.html
8.
▲
by
deusu
10y ago
Bookmarked. Thanks!
9.
▲
by
deusu
10y ago
File formats will be documented when I publish the data-files in a few weeks. What do you mean with postings? The main index is split into 32 shards (there is also an additional news-index which is updated about every 5-10 minutes). Each sh
10.
▲
by
deusu
10y ago
I have the filter implemented now. It's not perfect yet, but it already filters out a lot of the NSFW stuff. Unless you explicitly search for it. I'm gonna further improve this over the next days. Right now it's just a quick&
11.
▲
by
deusu
10y ago
I don't know sphinx at all, and my knowledge of lucene is very limited. Which means I don't know how they would compare to DeuSu.
12.
▲
by
deusu
10y ago
Thank you! Depending on who you are (there were 2 bitcoin donations today), you funded either about 18 or 28 hours of operations. :)
13.
▲
by
deusu
10y ago
The software is already open-source. A free search API will be fully available probably next week. It's in testing already. It's just a matter of putting the finishing touches on the documentation. And the crawl- and index-data wi
14.
▲
by
deusu
10y ago
Only ASCII and German umlauts (äöüß) at the moment. The parser needs rewriting. It was originally written in pre-unicode times. :)
15.
▲
by
deusu
10y ago
Originally it was written in Delphi. But I now use FreePascal for the development. I'm even compiling both Windows and Linux versions on my Linux machine.
16.
▲
by
deusu
10y ago
Yes, it would be better. The snippets are currently the first 255 characters of the page's text. For snippets to be customized to the search term, I would have to store all the text of the page. And that would require a lot more disk s
17.
▲
by
deusu
10y ago
Some issues that appeared over the years: Block outgoing connects to local IP nets in your firewall. Otherwise your hosting provider might think you are trying to hack them. Apparently there are a lot of links out there that point to hosts
18.
▲
by
deusu
10y ago
4 servers in total. 2 are used for crawling, index-building and raw-data storage. Quadcore, 32gb RAM, 4tb HDD and 1gbit/s internet connection on each of these. They are rented and in a big data-center. Crawling uses "only" ab
19.
▲
by
deusu
10y ago
It's all open-source. So, yes.
20.
▲
by
deusu
10y ago
I will publish the index for download in a few weeks. I'm currently working on the documentation. Oh, and I will publish the raw crawl-data too. Everything together is about 2.5tb. There is also a free API in beta-test right now. Will
21.
▲
by
deusu
10y ago
Thx. But all the traffic from here is currently driving the servers to their limit. Queries are already slowing down a bit because of imminent overload. Usually the average query takes about 250ms. Currently the average is at 334ms.
22.
▲
by
deusu
10y ago
And why should it? You are already at the destination. No need to find it. :)
23.
▲
by
deusu
10y ago
A fresh recrawl is currently running. Should take about 2-3 months. Newly crawled data will gradually replace older data during that time.
24.
▲
by
deusu
10y ago
In my experience this is usually caused by the fact that even 2bn pages aren't that many nowadays. The index needs to get bigger to better find (and rank) long-tail results like queries like this.
25.
▲
by
deusu
10y ago
I hadn't even thought about that. But it should be pretty easy to do in post-processing. I just have to take a list of "porn" keywords. If none of them occurs in the query, but in a search-result, then that result gets downra
26.
▲
Show HN: Open-source search engine with 2bn-page index
(deusu.org)
229 points
by
deusu
10y ago
|
139 comments
27.
▲
by
deusu
11y ago
I think Google once mentioned that each day a surprising amount of searches are unique. They were never done before. If memory serves me correct that number was 30-40%.
28.
▲
by
deusu
11y ago
I'm always open to new business opportunities. :) What would be more useful to you, the raw data - meaning for each page a list of the keywords on it - or the reverse-word-index? Raw-data may be better for batch-processing or running m
29.
▲
by
deusu
11y ago
(I'm not with CommonSearch. I have my own project that crawls extensively though.) You do realize that you are talking about potentially a LOT of data? To give you an example: The word "work" occurs on about 4% of all web-p
30.
▲
by
deusu
11y ago
Crawler works from a single IP. User-Agent is fixed to the robot's UA. Cookies are totally ignored. The search-engine works with just 2 servers. Crawler/Indexer and Webserver/queries. Crawler is a root-server with 1gbit/
More ›