Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
raphaelty
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
raphaelty
3y ago
It's because of the loss of the model. I ask the model to produce a higher similarity between the query and the positive document rather than between the query and the negative document. I'll add more losses soon so there are more
32.
▲
by
raphaelty
3y ago
Nice, it might already be compatible with BGE, I'll try it and add it to the documentation soon
33.
▲
by
raphaelty
3y ago
Yes exactly
34.
▲
by
raphaelty
3y ago
In the documentation there is an evaluation module with detailed informations. The idea is to gather relevant pairs of queries and documents that are not part of the training set. Then the idea is to measure, using various metrics, how your
35.
▲
by
raphaelty
3y ago
Hi, there is a single loss right now, but I plan to add some Sentence Transformers losses. ColBERT is slow as a retriever, but is quite efficient as a Ranker on GPU (way faster than cross-encoder). I plan to release pre-trained checkpoints
36.
▲
Show HN: ColBERT Build from Sentence Transformers
(github.com)
66 points
by
raphaelty
3y ago
|
18 comments
37.
▲
Show HN: Fine-Tuning Splade and SparseEmbed LMs
(github.com)
1 points
by
raphaelty
3y ago
|
0 comments
38.
▲
Cherche: Semantic Search
(github.com)
3 points
by
raphaelty
3y ago
|
0 comments
39.
▲
Minimalist semantic search with Cherche 2.0
(github.com)
3 points
by
raphaelty
3y ago
|
1 comments
40.
▲
by
raphaelty
3y ago
Cherche 2.0 is now available, and it's been optimized for batch-computing, along with other new features. Whether you're a practitioner, researcher, or hacker interested in semantic search, Cherche might be a good fit for your nee
41.
▲
by
raphaelty
3y ago
I build my personnal search engine which record things I like on twitter, blog posts etc.. It automatically calls those APIs using Github Action and store them in an open source database (json file) I actualy use it at least twice a week to
42.
▲
Python client for knowledge graph exploration
(github.com)
2 points
by
raphaelty
4y ago
|
0 comments
43.
▲
by
raphaelty
5y ago
I think 10 million documents is a large corpus. A retriever like Sklearn TfIdf will have a hard time handling it in a reasonable time. The main goal of Cherche is to prototype a neural search engine quickly and with a large choice of retrie
44.
▲
by
raphaelty
5y ago
Thank you for these great resources.
45.
▲
by
raphaelty
5y ago
Hi, 1) The dependency on the Elasticsearch python client allows Elasticsearch to be used as a retriever. The same goes for Lunr. It might be interesting to separate the different dependencies. 2) Of course I'll update it.
46.
▲
by
raphaelty
5y ago
I'm more used to reading than posting on Hacker News. I'll do better next time. :)
47.
▲
Neural search library in Python for medium-sized corpora
4 points
by
raphaelty
5y ago
|
4 comments
48.
▲
Neural Search for medium sized corpora
(github.com)
53 points
by
raphaelty
5y ago
|
7 comments
49.
▲
Knowledge Graphs Emb. With PyTorch
(github.com)
3 points
by
raphaelty
6y ago
|
1 comments
50.
▲
by
raphaelty
6y ago
Knowledges graphs are structured resources in the form of graphs that contain knowledge. These resources are used in a large number of applications linked to the machine learning. I just published a library dedicated to knowledges graphs em