4 ms·
I've been meaning to attempt running a custom search engine for particular sites I've 'bookmarked'. Some sites contain gold that could be useful in the future a
by b1zguy 3y ago
I've been meaning to attempt running a custom search engine for particular sites I've 'bookmarked'. Some sites contain gold that could be useful in the future and is not often discovered in Google results.
Should I go the Postgres/Elasticsearch route or are somewhat out-of-the-box solutions available?
- Alifatisk 3y agoSearchkick gem + Elasticsearch is a good combo
- sudobash1 3y agoI am wanting to do something similar. Archivebox seems to be the best solution for this sort of self-hosted, searchable web archive. It has multiple search back-ends and plugins to sync browser bookmarks (or even history). I haven't finished getting it set up though, so take this recommendation with a hefty grain of salt.
- SpriglyElixir12 3y agoHow would something like this work in practice? Would you generate any tags or summaries per site when inserting it into the db?
- janalsncm 3y agoYou could run a full text search or search against an auto-generated summary. Or if you want to be fancy, use semantic search like in Retrieval Augmented Generation.
- sudobash1 3y agoArchiveBox can extract text from HTML (and possibly PDFs too). I think it can be configured to extract subtitles from YouTube videos as well. So it can do full text searches. Basically you could have your own, offline & curated search-engine.
- bshipp 3y agoFor such a light demand and fixed site requirements, a single-file sqlite dB is probably best. Modern Sqlite has full-text capabilities that are quite powerful and relatively easy to implement. https://www.sqlite.org/fts5.html https://www.sqlite.org/fts5.html
- tudorg 3y agoDo you prefer it locally or in the cloud? If in the cloud, check out Xata (the domain of the blog post here).
- binarymax 3y agoFor something small with a minimal footprint, I'd recommend Typesense. https://github.com/typesense/typesense https://github.com/typesense/typesense Elasticsearch is heavy, and relational databases with search bolted on (like Postgres or SQLite) aren't great.
- bshipp 3y agoIt depends on what the user requirements are. FTS works pretty well with both Postgres and SQLite, in my experience. Here's a git repo someone can modify to do a cross comparison on a specific dataset, if they are interested. It doesn't seem to indicate the RMDBs are outclassed in a small-scale FTS implementation. https://github.com/VADOSWARE/fts-benchmark https://github.com/VADOSWARE/fts-benchmark
- binarymax 3y agoFor personal use nobody cares about 100ms vs 10ms response. What they do care about is relevance. Consider the following from those repo outputs: Typesense [timing] phrase [superman]: returned [28] results in 4.222797.ms [timing] phrase [suprman]: returned [28] results in 3.663458.ms SQLite [timing] phrase [superman]: returned [47] results in 0.351138.ms [timing] phrase [suprman]: returned [0] results in 0.07513.ms So SQLite is faster, but who cares? I want things like relevance and typo resilience without having to configure anything.
- bshipp 3y agoI'm not trying to be argumentative. As long as people find a solution they're happy with, I think that's great. For me, I'm far less interested in handling typos, but I can see how it would be valuable in many applications. I'm usually less interested in tying in and learning another set of services if I can get 90% of the way there with one, but leaving the option of adding it later if additional requirements make it necessary.
- evdubs 3y agoThe article covers typo resilience in the section "Typo tolerance / fuzzy search". This adds a step between query entry and text search where you find the similarity of query words to unique lexemes if the word is not a lexeme. Seems like a reasonable compromise to me?
- b1zguy 3y agoEdit: I forgot to add how would I add the webpage to the databases already suggested here? Do I need to use a separate program to spider/index each site, and check for its updates?
- bshipp 3y agoIf you're looking for a turn-key solution, I'd have to dig a little. I generally write a scraper in python that dumps into a database or flat file (depending on number of records I'm hunting). Scraping is a separate subject, but once you write one you can generally reuse relevant portions for many others. If you can get adept at a scraping framework like Scrapy you can do it fairly quickly, but there aren't many tools that work out of the box for every site you'll encounter. Once you've written the spider, it's generally able to be rerun for updates unless the site code is dramatically altered. It really comes down to how brittle the spider is coded (i.e. hunting for specific heading sizes or fonts or something) instead of grabbing the underlying JSON/XHR that doesn't usually change frequently. 1. https://scrapy.org https://scrapy.org
- busymom0 3y agoDepending upon the type of content, one might want to look into using the Readability (Browder's reader view) to parse the webpage. It will give you all the useful info without the junk. Then you can put it in the DB as needed. https://github.com/mozilla/readability https://github.com/mozilla/readability Btw, readability, is also available in few other languages like Kotlin: https://github.com/dankito/Readability4J https://github.com/dankito/Readability4J