14 ms·
Show HN: Full text search Project Gutenberg (60m paragraphs)
- tornato7 6y agoI'd love to see this dataset used as a performance and relevance benchmark for different search engines!
- gutensearch 6y agoThat was definitely part of the original plan! I spotted two other attempts [1] [2] here using BERT and ElasticSearch respectively. The main performance issue with the Postgres FTS approach (possibly also the others?) is ranking. Matching results uses the index, but ts_rank cannot. Most of the time, few results are returned and the front end gets its answer in ~300ms including formatting the text for the front end (~20ms without). However, a reasonably common sentence will return tens or hundreds of thousands of rows, which takes a minute or more to get ranked. In production, this could be worked around by tracking and caching such queries if they are common enough. I'd love to hear from anyone experienced with the other options (Lucene, Solr, ElasticSearch, etc.) whether and how they get around this. [1] https://news.ycombinator.com/item?id=19095963 https://news.ycombinator.com/item?id=19095963 [2] https://news.ycombinator.com/item?id=6562126 https://news.ycombinator.com/item?id=6562126 (the link does not load for me)
- ngrilly 6y agoI suggest to have a look at https://github.com/postgrespro/rum https://github.com/postgrespro/rum if you haven’t yet. It solves the issue of slow ranking in PostgreSQL FTS.
- karterk 6y agoWhat kind of hardware are you using to host the Postgres instance?
- gutensearch 6y agoSame place as the app: a Start-2-M-SSD from online.net in their AMS1 DC (Amsterdam). Subset of sudo lshw --short: Class Description ====================================================== processor Intel(R) Atom(TM) CPU C2750 @ 2.40GHz memory 16GiB System Memory disk 256GB Micron_1100_MTFD
- kristopolous 6y ago"What is a cynic" the famous line from Lady Windmere's Fan comes up empty. You can find it here: https://www.gutenberg.org/files/790/790-h/790-h.htm https://www.gutenberg.org/files/790/790-h/790-h.htm The next line I can find though. How are you parsing?
- gutensearch 6y agoThanks for noticing! To be more specific, "what is a" are stop words and "cynic" is very common, so a lot of rows are returned (see my other comment). ts_rank takes too long to rank them, and the server times out, leaving you with the previous query's table because I didn't take the time to program a correct response to this issue. "Cecil Graham. What is a cynic?" returns Lady Windermere's Fan almost instantly. The workarounds I've thought of would be to cache these queries (assuming I've seen them before, and after I've set up logging), buy a larger server, or pay Second Quadrant to speed up ts_rank... I'd love any suggestions from more experienced Postgres engineers!
- gutensearch 6y agoThanks for noticing! To be more specific, "what is a" are stop words and "cynic" is very common, so a lot of rows are returned (see my other comment). ts_rank takes too long to rank them, and the server times out, leaving you with the previous query's table because I didn't take the time to program a correct response to this issue. "Cecil Graham. What is a cynic?" returns Lady Windermere's Fan almost instantly. The workarounds I've thought of would be to cache these queries (assuming I've seen them before, and after I've set up logging), buy a larger server, or pay Second Quadrant to speed up ts_rank... I'd love any suggestions from more experienced Postgres engineers! Edit to your edit, re parsing. The subset of rows returned follows: where language = %s::regconfig and textsearchable_index_col @@ phraseto_tsquery(%s::regconfig, %s) and relevance is determined by: ts_rank_cd(textsearchable_index_col, phraseto_tsquery(%s::regconfig, %s), 32) with the %s being language and paragraph respectively.
- kristopolous 6y agofull-text searching is contentious, it's a rabbit-hole of holy-wars. I don't believe I'm going to summon cranks with the following: There's more specialized tools for your usecase than postgres and you should be looking into "n-gram indexing". Lucene based systems such as elasticsearch are quite popular there's also sphinx and xapian, also fairly widespread. You need to read the documentation and configure them, they are flexible and need to be tuned to this usecase. In the end, there is no "correct" way to do things. For instance, sometimes stemming words is the right way to go, but if you are say doing a medical system where two very different medicines could be spelled similar and stem to the same word, a mistake could lead to death, so no, stemming is very bad here. Sometimes stop words is the way to go, but if people are looking for quotes, such as "to be or not to be" well now you have the empty string, splendid. So yeah, configure configure configure. This may bring out the cranks: If you want to roll-your-own nosql systems like redis and mongo or couch seem to work really well (I've rolled my own in all 3 on separate occasions). I guarantee there's advanced features of maria and postgres that aren't widely used and some people reading this will confidently claim they are superior but I assure you, that is a minority opinion. Most people go with the other options. If you ever doubt it, ask the actual developers of postgres or maria on chat. They are extremely nice and way more honest about the limitations then their hardcore fans are. The databases are under constant development with advanced features and you'll learn a lot (really, they are both super chill dev communities, impressively so). Perhaps your solution (as mine has been) is a hybrid. You can store the existence for instance, in one system and the offsets and chunks in another so you get a parallelizable asynchronous worker pipeline, it's impressively fast when you horizontally scale it. <0.5 sec for multiple terabytes of text (and yes, I'm talking nvme/hundreds of gb of RAM per node/>=10gb network). I've legitimately just done random queries to marvel at the speed I really hope I save myself from the rock throwing.
- gerardnll 6y agoI cannot choose another language, some weird css/html issue happens that hides all other fields. Looks like an overflow:hidden hides the possible selections.
- gutensearch 6y agoThanks for the bug report! It has to do (I think) with Dash's columnar layout which unfolds the menu over the next few columns at least in Chrome. The quick workaround I found was to type out the language until it appears below and click on it or finish typing, then press enter. This should select it. I'd love to hear from other Dash developers who've had, and solved this issue.
- tkgally 6y agoThis is a very useful tool. I’ve known people doing research on the evolution of grammar, vocabulary, literary style, etc. who use only small subsets of the Project Gutenberg data. I'm sure they would appreciate being able to search the entire corpus. The corpus-search functions those researchers use include wildcards, exact-phrase specification with quotation marks, proximity searches, and Boolean search strings. When you have a chance, you might want to add a list of the syntax formats that currently work. (I tried using * as a wildcard in a phrase surrounded by quotation marks, and it didn’t seem to work.) One small improvement you could make would be to widen the “Search terms” field so that longer search strings are visible.
- gutensearch 6y agoThank you for taking the time to lay out feature requests in such details! I really appreciate it. The current search box is a wrapper around Postgres phraseto_tsquery [1] whilst the Discovery tab uses plainto_tsquery, so you could play with either as an ersatz for some of these features for now, although special characters might get stripped or parsed incorrectly. Do you know where the people you are talking about hang out online (for example, subreddits)? I'd love to get in touch with them once the features are built and for more general feedback. [1] https://www.postgresql.org/docs/12/textsearch-controls.html https://www.postgresql.org/docs/12/textsearch-controls.html
- iiv 6y agoI know /r/CompLing on Reddit is quite popular.
- StavrosK 6y agoI'm curious if you ever tried MeiliSearch for this, I tried it recently for something unrelated and had a very good experience with it. Since you have the corpus already, it might be worth trying and seeing if it speeds search up? I'd be interested in the results either way.
- gutensearch 6y ago
- Kaknut 6y agoThis is indeed very useful. But I think there are some bugs like when I click on dropdown to choose language it breaks (malfunction)
- gutensearch 6y agoThanks! There's a workaround to the dropdown issue, see this other comment thread: https://news.ycombinator.com/item?id=25890458 https://news.ycombinator.com/item?id=25890458
- snikolaev 6y agoHi. Nice project! I don't know why there's no BM25 (or at least SOME TF-IDF implementation) in Postgres FTS, but if you decide you need it (and more languages support and highlighting and lower response time) ping us at contact@manticoresearch.com and we'll help you with integrating your postgres dataset with Manticore Search. 60M docs shouldn't be a problem at all (should take about an hour to make an index) and you'll get proper ranking and nice highlighting with just few lines of new code. Here's an interactive course about indexing from mysql https://play.manticoresearch.com/mysql/ https://play.manticoresearch.com/mysql/ , but with postgres it's the same.
- gutensearch 6y agoThank you for your generous offer of help! I look forward to taking it up (may take a while as I'm about to move countries and quarantine). In particular I love that one of the examples in your comment history is in Latin as that language is not currently supported by Postgres FTS. Are Latin and Ancient Greek supported by Manticore? (dare I hope for Anglo Saxon...)
- snikolaev 6y agoIn terms of advanced NLP (stemming, lemmatization, stopwords, wordforms) - no. In terms of just general tokenization - I've never dealt with Latin and Ancient Greek characters (if there're specific characters for those languages), but if even they are not supported by default it's not a problem to add them in config (https://mnt.cr/charset_table https://mnt.cr/charset_table)
- yorwba 6y agoFor the character mappings, it might be useful to have a look at the config for https://tatoeba.org https://tatoeba.org (or rather, the PHP script that generates the config): https://github.com/Tatoeba/tatoeba2/blob/dev/src/Shell/SphinxConfShell.php https://github.com/Tatoeba/tatoeba2/blob/dev/src/Shell/Sphin... There's one big list of mappings for almost every script under the sun, including Greek. (With mappings like 'U+1F08..U+1F0F->U+1F00..U+1F07' turning U+1F08 Ἀ [CAPITAL ALPHA WITH PSILI] into U+1F00 ἀ [SMALL ALPHA WITH PSILI], and the same for seven other accented alphas. I've considered turning them all into unaccented alpha instead, but I don't know enough about Greek orthography to decide that.) https://github.com/Tatoeba/tatoeba2/blob/3170f7326ad2939c691ba23f05edc65026ade355/src/Shell/SphinxConfShell.php#L123 https://github.com/Tatoeba/tatoeba2/blob/3170f7326ad2939c691... For Latin, there are some special exceptions so that "GAIVS IVLIVS CAESAR" and "Gaius Julius Caesar" are treated the same: https://github.com/Tatoeba/tatoeba2/blob/3170f7326ad2939c691ba23f05edc65026ade355/src/Shell/SphinxConfShell.php#L322 https://github.com/Tatoeba/tatoeba2/blob/3170f7326ad2939c691... It's not beautiful, but it's used in production. People who don't need to support quite as many languages as Tatoeba will probably want a simpler config, but it might still be useful as a reference.
- JensRantil 6y agoNice project! You might want to ask a designer to give it a look. UI is a little cluttered.
- gutensearch 6y agoThanks! I would love for an experienced designer to join the project (and DevOps, and front end, and...). My email is in the profile.
- gebt 6y agoWhat useful thing! It would be much better if it was possible to pass search query using GET requests. e.g. https://gutensearch.com/?q=avicenna https://gutensearch.com/?q=avicenna
- gutensearch 6y agoAn API is already on my roadmap! I couldn't quite figure the state of passing query parameters in the URL with Dash, though.
- gebt 6y agoHaving an API is good but not suitable for what I mean. I'm not familiar with Dash, but IN WORST CASE I think you can add a route to the nginx (an additional app) to pass GET parameters to the app.
- _Microft 6y agoOh, so Project Gutenberg is still a thing? I used to use it until they blocked German users (context: they were asked to prevent access to certain pieces for German users but they decided to go nuclear instead). Nowadays I just go to LibGen when I want to have a look into a book. That LibGen doesn’t limit itself to works in the public domain is rather a feature ;)
- gutensearch 6y agoAnd this is why the server is in Amsterdam, even though I have had good experiences with Hetzner in the past. I was quite sad to read about the case in 2018, and it is unfortunate that it is still not resolved.
- SahAssar 6y agoHetzner has locations in finland too if that would work.
- IndySun 6y agoWhat's the German story?
- robin_reala 6y agoWork that’s PD in the US but not in Germany. German rights-holders took PG to court (or rather, a poor German sysadmin for PG), court demanded PG pull the titles for Germany, PG refused saying that it was down to the users to determine PD status for their home country, and because of HTTPS the German court’s only option was to block PG at the domain level.
- _Microft 6y agoProject Gutenberg themselves decided to block all access from Germany instead of just the items in question. „On February 9 2018, the Court issued a judgment granting essentially most of the Plaintiff's demands. The Court did not order that the 18 items no longer be made available by Project Gutenberg, and instead wrote that it is sufficient to instead make them no longer accessible to German Internet (IP) addresses. PGLAF complied with the Court's order on February 28, 2018 by blocking all access to www.gutenberg.org and sub-pages to all of Germany.“, from https://cand.pglaf.org/germany/index.html https://cand.pglaf.org/germany/index.html
- jjt-yn_t 6y agoI changed the default search to 'wherefore' and it came up with results for 'it was the best of times' which existed upon loading the web page. Twice. So a bug somewhere.
- imaginenore 6y agoI searched for "it was a dark and stormy night", and it didn't find the origin (1830 "Paul Clifford" novel).
- nestorD 6y agoGreat ! It is not far from one of my dream project : numerising a maximum of old text and let historian do research on them using state of the art tools (that work across synonims and languages) with parameters to restrict by time of publication and localisation obviously.
- toomuchtodo 6y agoHave you considered working with the Internet Archive on this across their corpus? They are open to such work being done. And if some of the material you need isn’t in the archive, let’s get it in there.
- nestorD 6y agoI have not but I am going to file the idea, it would indeed be a good starting point.
- loughnane 6y agoI love this. So often when I search for passages I get bombarded with links to quote websites. This is much closer to what I’m looking for. EDIT: and it’s going to be open sourced? I love it
- gutensearch 6y agoThanks! I had the exact same problem and eventually it got me to do something about it. It is particularly bad with writers from antiquity or with a lot of popular appeal. I've begun adding to this repository, it'll come in piece by piece as I clean up the code: https://github.com/cordb/gutensearch https://github.com/cordb/gutensearch
- abhayhegde 6y agoSeems like a nice project. However, I am currently receiving PR_CONNECT_RESET_ERROR on Firefox. Bringing it to your attention.
- gutensearch 6y agoInteresting! I don't recall seeing this before, and the app otherwise held up to the HN hug of death which was unexpected given nothing is optimised. Can you recall the steps that led to the error?
- abhayhegde 6y agoI was not able to access it until I last checked on yesterday night. I was very interested in implementing this project myself, but did not have enough SQL skills to go about it, and hence I was naturally curious. It was working for a brief period where I accessed it, however seems to be down for me again. It is an amazing site for sure. I am very happy one can just search for any phrase and it shows up with results in a matter of seconds! Allowing fuzzier results is a very useful feature I must say. The site may use a few improvements on CSS part as you may have already noticed (search boxes get smaller while typing, including dark theme, cleaning up the results table, etc). Also, "Discovery" section was not displaying the correct results although the query was executed for some time (I searched "days of our lives" in place of "call me Ismael"). With the same parameters, the search engine shows results anyway, so not a big deal. Please release the pg_dump some day. It would be extremely useful.
- gutensearch 6y agoThanks for the details! I tried your search string but could not replicate the exact bug (either app not working at all, or your error message). What is your setup and connection speed? Please feel free to email me instead at contact@ if you would prefer to preserve your privacy. "Days of our lives" returns fine on the Search tab and times out (as I assumed it would, with such common words) on the Discovery tab because the broader plainto_tsquery is returning too many rows to rank in time. ...which made me think, Discovery ranks by random(), so I do not need to calculate ts_rank... time to push a quick fix! Regarding pg_dump, I originally wanted to do this, but at 60GB a pop the bandwidth would be quite expensive if the thing got popular at all; I recommend you head to the repository [1] and build it locally instead. [1] https://github.com/cordb/gutensearch https://github.com/cordb/gutensearch
- DerWOK 6y agoHm. I cannot do language selection on iOS Safari. Even not with the mentioned workaround of first typing the language name prefix.
- Ninjinka 6y agoI enjoy being able to grep the 200,000 books I have downloaded as part of the Pile that was on HN a while back. It allowed me to do things like "show me all the paragraphs where 'Edmund Burke' and 'patriotism' appear close to each other." This seems to be a similar thing, just with less books.
- rahimnathwani 6y agoThis is really cool. Something like this should exist. It seems like you could do it more easily, include all recent additions, and have faster search responses: 1. Mirror the current gutenberg archive (e.g. rsync -av --del aleph.gutenberg.org::gutenberg gutenberg) 2. Install recoll-webui from https://www.lesbonscomptes.com/recoll/pages/recoll-webui-install-wsgi.html https://www.lesbonscomptes.com/recoll/pages/recoll-webui-ins... or using docker-recoll-webui: https://github.com/sunde41/recoll https://github.com/sunde41/recoll 3. Run the recoll indexer 4. Each week, repeat steps #1 and #3
- ilaksh 6y agoI am sure the Postgres full text search is pretty good by now. But I always expect projects like these to eventually reluctantly switch over to Elasticsearch in order to handle 10-20% of the queries better.