4 ms·
We also built a similar semantic search engine for retrieving open-source projects. We experimented with HyDE and guessing user queries based on a corpus, both
by redskyluan 3y ago
We also built a similar semantic search engine for retrieving open-source projects. We experimented with HyDE and guessing user queries based on a corpus, both of which were very successful.
For embeddings, we chose BGE, but found they seemed to overfit on Beir. CohereV3 and Voyage appeared to perform better in practice.
For retrieval, we used OpenSearch + ZillizCloud Serverless + Cohere ranking, providing us with maximum flexibility and search effectiveness.
To be mentioned, build a evaluation dataset is important. This helps me to improve the quality.