3 ms·
I thought text search was always the first thing you try, then fuzzy search, then you go for RAG
by lacedeconstruct 1mo ago
I thought text search was always the first thing you try, then fuzzy search, then you go for RAG
- ozim 1mo agoI think Bitwarden implemented some vector search in their password search feature ... totally annoying it gives me back all kinds of stuff that I don't care. I want fuzzy search like 95% of time and then I might consider having additional list of things that can be suggested by vector search.
- a1o 1mo agoA good UI could do these and also exact match, give some point system to the results, then order them and perhaps use a bold highlight to reflect what parts of the input query reflected in each result.
- gwerbin 1mo agoBandcamp has had legendarily bad semantic search for as long as they've been around. It's often completely impossible to find an artist or album or song even when you type the exact name.
- t_mahmood 1mo agoahh now I realize why I get so much completely irrelevant search results in many sites recently. I mean I'm searching for betel and you're giving me nuts. haha
- itintheory 1mo agoBitwarden has lost the plot. The most recent Windows update is so bad. It has way lower information density in the UI, more buttons to click for the same use, no longer puts focus on the search field by default (this one makes me irrationally angry), and on one of my Win 11 installs can't lock the vault, manually or automatically. How could they mess up such a simple app that worked fine for so long?! What perverse incentives caused this nonsense?!
- wongarsu 1mo agoIt's not like a simple embedding search takes that much longer to implement. Especially on short descriptions where you don't have to deal with chunking. And if you let an LLM write the code it's even less of a difference. Combine that with embedding search promising to solve all your search problems, and I understand why people often skip over full text search and go straight to embeddings
- j0selit0 1mo agoI wish everyone thought like you, in my experience unfortunately it's not the case
- EagnaIonat 1mo agoEven that is an oversimplification unless you are doing something very basic. Volume of documents, size of documents, versioning, frequency of update, documents similar or overlapping information, how much or exactly what you need for the LLM to understand, AI friendly documents, who has access and at what level, blue teaming, red teaming, multi-lingual, does the LLM know the domain language of the user and documents. I probably missed a few things even with that.