Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Kerollmops
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
Heed v0.20: Safest and most maintained Rust wrapper for the LMDB key-value store
(github.com)
3 points
by
Kerollmops
2y ago
|
0 comments
32.
▲
Meilisearch Updates a Millions Vector Embeddings Database in Under a Minute
(blog.kerollmops.com)
9 points
by
Kerollmops
3y ago
|
0 comments
33.
▲
by
Kerollmops
3y ago
Indeed, there could probably be cases where you see higher deserialization cost than raw lists of integers. But when it comes to high number of integers I can confirm that it is much more efficient.
34.
▲
by
Kerollmops
3y ago
> Interesting. Could you elaborate on the benefit of this? I don't know on what I can elaborate. Storing integers that are near each other is much more optimal in a RoaringBitmap than in a flat array. The reason is that it will only
35.
▲
by
Kerollmops
3y ago
You can use arroy, the vector store, directly in your library. It is in Rust and runs on top of LMDB which manages a disk file.
36.
▲
by
Kerollmops
3y ago
And when you'll use the v1.6 you see a much better indexation speed. https://x.com/kerollmops/status/1734576622303404317
37.
▲
by
Kerollmops
3y ago
They recently released a new YouTube stats website and they use Meilisearch in the frontend. https://news.ycombinator.com/item?id=38681328
38.
▲
by
Kerollmops
3y ago
Thank you for the feedback. You should probably look at our blog post explaining how to index in smarter way [1]. You should probably come and talk to us on Discord. We work with MrBeast and they don't have any issue indexing millions
39.
▲
by
Kerollmops
3y ago
Indeed, you are right about this. This is why we will release geo-replication on the cloud in Q1 next year. We heard you and worked hard on both of those features. We already have a great working proof-of-concept.
40.
▲
by
Kerollmops
3y ago
It's my own company and I work hard on both Meilisearch, the keyword search part since 2018 and the soon-to-be-released semantic search part. The hybrid search will be also part of the v1.6 release in january. I took time to write thos
41.
▲
Meilisearch expands search power with Arroy's filtered disk ANN
(blog.kerollmops.com)
75 points
by
Kerollmops
3y ago
|
26 comments
42.
▲
Spotify-Inspired: Elevating Meilisearch with Hybrid Search and Rust
(blog.kerollmops.com)
2 points
by
Kerollmops
3y ago
|
0 comments
43.
▲
Arroy: Approximate Nearest Neighbors in Rust and optimized for memory usage
(github.com)
5 points
by
Kerollmops
3y ago
|
0 comments
44.
▲
Meilisearch Across the Semantic Verse
(github.com)
5 points
by
Kerollmops
3y ago
|
0 comments
45.
▲
by
Kerollmops
3y ago
That looks pretty cool. Would you mind explaining a little bit how you did that into more detail? Like, you ask Meilisearch some documents (keyword search) and then ask the LLMs score of those and refine the list a bit?
46.
▲
by
Kerollmops
3y ago
> Yup, that's exactly what I'm saying there is something wrong with data model. We are an Open Source project and know we can do better on the indexing speed. We already did much work on that subject. We enabled the auto-batchi
47.
▲
by
Kerollmops
3y ago
Hey rkwasny, Which version of Meilisearch were you using? We chose to use LMDB instead of RocksDB because it is a much faster, memory and CPU-efficient key-value store. However, it would probably be much quicker to insert all those inverted
48.
▲
by
Kerollmops
3y ago
Co-founder and Tech Lead here, Thanks to Louis' work and the design of LMDB, we can efficiently use the virtual address space of the OS to let it manage in the best way the memory Meilisearch is using. We continuously improve the index
49.
▲
by
Kerollmops
4y ago
We are not in contact with beir or the owner of the bei-cellar oganisation. However, we started tracking our relevancy with the TREC 4 & TREC 5 data which are provided by the NIST organisation [1]. I can only tell that the results are v
50.
▲
by
Kerollmops
4y ago
What’s funny is that (1) doesn’t look like a real limit when you know that the first Harry Potter book is nearly 77000 words. The recommended way is to split your documents by paragraph to increase relevancy, this way you can see the exact
51.
▲
by
Kerollmops
4y ago
Thank you very much for this amazing feedback, really appreciated. We did a lot of improvement to the indexing part of the engine and now can auto-batch updates which gaves incredible improvements. We will continue to work on this in 2023.
52.
▲
by
Kerollmops
4y ago
I am the co-founder and maintainer of the engine, and I confirm we have some localized unsafe blocks for when we interface with the C library: LMDB. However, I prefer having a few unsafe blocks that I can review carefully than a single one
53.
▲
by
Kerollmops
5y ago
A previous version of MeiliSearch was using RocksDB and we were having a lot of trouble using it, a lot of setup to do to make sure that we were not killed by the OS due to OOM or even fixing a lot of strange segfault by patching the RocksD
54.
▲
by
Kerollmops
5y ago
As we don’t use RocksDB but LMDB, we use a lot less real memory than key-value stores that uses a user-side cache system. LMDB is memory mapped and therefore let the OS manage memory for it. Typesense uses RocksDB and ElasticSearch a custom
55.
▲
by
Kerollmops
5y ago
Hey, this no more a prototype, it is the internal engine under MeiliSearch. I forgot to update the README :)
56.
▲
by
Kerollmops
6y ago
Nice work, it’s pretty awesome to be able to plug a search engine as easily as adding a plugin. I just have one question about the recently published benchmarks, are the latency based on the two-word search "hello world"?
57.
▲
by
Kerollmops
7y ago
I would say that you must add your own nginx (or else) in front of our HTTP only engine, in term of fault tolerance we are working on high availability.
58.
▲
by
Kerollmops
7y ago
To make MeiliSearch expose the documents that are stored in your PostgreSQL (or any other database) you must extract them and store them in our engine using the HTTP API we provide to you. https://docs.meilisearch.com/refere
59.
▲
by
Kerollmops
7y ago
This is not something that MeiliSearch supports currently but I am working on making the engine be able to index other formats than JSON, I saw great performance improvements when indexing simple CSVs. We will probably make MeiliSearch acce
60.
▲
by
Kerollmops
7y ago
Just to add a little note here, we are currently working on the functionality of multi-filter queries, because we are aware of our community!
More ›