3 ms·
Can you elaborate? We have technicians searching in different languages. Also our knowledge base is often in different languages. I just don't see how full text
by kaon_2 1mo ago
Can you elaborate? We have technicians searching in different languages. Also our knowledge base is often in different languages. I just don't see how full text search can work? Maybe in a problem space like a wiki where people always know what to search for?
- jon-wood 1mo agoInstinctively this feels like a two phase problem - start with some machine translation into a single spoken language and index that, then when people are querying do the same thing. When returning search results show them in the original language.
- whilenot-dev 1mo agoWhy not create indexes for multiple languages, as that would also avoid double translation issues (e.g. GER [query] → ENG [index] → GER [document])?
- j0selit0 1mo agoyou would also need to maintain multiple indexes in multiple languages. I never had to do that - but I assume it's a pain
- whilenot-dev 1mo agoIt's a matter of running a for-loop. You'd get faster response times (no query→index translation), but the storage requirements for the indexes would be larger.
- kaon_2 1mo agoYes we've tried. It works. But jargon is hard. RAG with embeddings works all the same. The LLM doesn't mind receiving sources in Italian, french and German, and then outputting the answer in Japanese while providing the verbatim German jargon term in brackets
- jameshart 1mo agoEmbedding search is effectively machine translation into a single common ‘language’ - embedding space - and then searching that; cleaner and less lossy than translating everything into English for searching, but harder to debug when it goes wrong.
- tantalor 1mo agoFTS like Elasticsearch supports cross-language (also called multi-language) search.
- hnfong 1mo agoYes. Thank you for pointing this out. I think there needs to be a linguist version of "what every programmer needs to know about (full?) text search"... I'm not a linguist and I don't study languages, but I know enough to realize if a text search system is not designed for a particular language, it simply won't work. (As an example, to implement English search in a system for a hobby project, I had to import a US/UK spelling wordlist, and implement the Porter Stemming Algorithm. This is just for "one" language, and probably does not cover the other "English" dialects. Imagine doing a different workaround for every language in existence...) RAG is actually a very language-agnostic way to work around those issues.
- usernametaken29 1mo agoI am a linguist and can say your view is too simplistic. This works only for things that are common enough to be mapped onto the same vector from the training set. Try searching for my company XY and it won’t work, because your embedding doesn’t have an adequate embedding, so you will end up having to build a translation table anyways..