Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fulmicoton
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
fulmicoton
3y ago
I had a look at this passage of the talk. Generally S3 is more obvious for OLAP where insert are very batchy and reads are large. For OLTP, the equation is less obvious. I'd be scared to suffer from the cost associated with PUT/G
32.
▲
by
fulmicoton
3y ago
Quickwit is compatible with S3-compatible object storages yes. I don't remember the feedback for R2 specifically, but we have users running Quickwit on S3, GFS, Azure, MinIO, Garage, IBM, and all of the major chinese clouds. > Secon
33.
▲
by
fulmicoton
3y ago
Zookeeper is not used by Quickwit at all. It is used by Kafka however.
34.
▲
by
fulmicoton
3y ago
Not yet. Only prefixes. Also you could probably cook something with an ngram tokenizer. Is it for a field with a high cardinality? If you tell us more about your use case, maybe we can find a workaround.
35.
▲
by
fulmicoton
3y ago
Web crawl can be an ok use case actually. The idea then would be to "reindex" the world. It might seem ludicrous, but to give you an idea, indexing CommonCrawl takes about a day with 8vCPUs.
36.
▲
by
fulmicoton
3y ago
Most users do not use Kafka/Zookeeper. The only external service for them is a S3 bucket. They then use the PushAPI. It is perfectly fine if you have only a couple of TB a day. For the crazy large use cases, like the one described in t
37.
▲
by
fulmicoton
3y ago
Quickwit is FTS too though. I think the difference comes from the fact they stored stuff on EBS while Quickwit stores its index on S3.
38.
▲
by
fulmicoton
3y ago
You can use AGPL for commercial use. The limitation is that you will have to opensource any patch you make to it.
39.
▲
by
fulmicoton
3y ago
It is a indexing vs search trade-off. Let's just consider IO, as it is the main effect here. With a columnar, you have to read all of the fields targeted by your query. If this is log, let's assume - 200B per line of logs - 100TB
40.
▲
by
fulmicoton
3y ago
I didn't know about Azure Premium blog storage! The pricing is very similar too.
41.
▲
by
fulmicoton
3y ago
I ran the benchmark at Quickwit. I confirm it works as intended. I was extremely excited about this feature, primarily interested in the decreased GET request cost, and secondly the lower latency. Unfortunately the price model puts it in a
42.
▲
by
fulmicoton
3y ago
> The other is that it implies serialised processing - you can't process anything > in parallel as you have single index threshold that defines what has been and > what has yet not been processed. Fortunately Kafka is partitio
43.
▲
by
fulmicoton
3y ago
That's a good list.
44.
▲
by
fulmicoton
3y ago
It has information from 2021. ChatGPT presents Quickwit as follows: As of my last knowledge update in September 2021, Quickwit is an open-source search engine infrastructure that is designed for building and deploying search solutions quick
45.
▲
by
fulmicoton
3y ago
Haha I regenerated the results something like last week, and the PISA engine is not in the list anymore :)
46.
▲
Why we tweaked Roaring Bitmaps
(quickwit.io)
3 points
by
fulmicoton
3y ago
|
0 comments
47.
▲
by
fulmicoton
3y ago
This is indeed correct. Yes, we have comparatively few actors. One thing that is not explained in the blog post however, is that we can run several indexing pipeline on the same machine. We have a bunch of tricks in our framework to run ~10
48.
▲
The Organic Software Manifesto
(quickwit.io)
2 points
by
fulmicoton
4y ago
|
0 comments
49.
▲
by
fulmicoton
4y ago
The trouble with spawn_blocking is that it runs on the tokio thread pool... By default the size of that thread pool is gigantic because it is meant to run blocking IO. For computation, you would want to have a number close to the number of
50.
▲
by
fulmicoton
4y ago
It can be used to search other things than logs, but it has to be large datasets of Append only data. Emails, Chat, Web Crawl data, logs...
51.
▲
by
fulmicoton
4y ago
And for those who cannot wait, there is quickwit :-) https://quickwit.io
52.
▲
by
fulmicoton
4y ago
pure algorithm & perf stuff... - exponential unrolled linked list: I don't know what they are called, so I called them that way https://fulmicoton.com/posts/tantivy-stacker/ . - radix heap. I actually had
53.
▲
by
fulmicoton
4y ago
In the beginning of the 90s I'd say there was a minitel in maybe 30% of homes in France (maybe a computer in <5%). That was the preferred to look up a phone number / address, and checking exam results.
54.
▲
by
fulmicoton
4y ago
This is correct yes. Rumor-mongering is more efficient both in term of detecting/propagating a node failure and in term of network overhead. We stopped using SWIM because it is too hard to get right, and because scuttlebutt allows all
55.
▲
by
fulmicoton
5y ago
Quickwit CEO here. For people interested, we are focusing on immutable data (logs of all kinds typically). https://github.com/quickwit-oss/quickwit We make it much 10x cheaper and easier to scale and operate. We store
56.
▲
by
fulmicoton
5y ago
With inlining maybe?
57.
▲
by
fulmicoton
5y ago
I agree! ... But I do wish I had some knobs in stable to hint the compiler in rust. Mark a condition as unpredictable, a branch as likely, a loop as likely to be long, etc. Some of it exists as intrinsics but this is not accessible in stabl
58.
▲
by
fulmicoton
5y ago
Not really. Lingua Franca is not really related to French.
59.
▲
by
fulmicoton
5y ago
A normal search experience (displaying a 20 hits search page) requires num segments * (1 + num terms * 2) + 20 GET requests. We have 180 segments for our commoncrawl index. So we can consider a generous upper bound of 1000 requests. The GE
60.
▲
by
fulmicoton
5y ago
There are only 180 splits. For this demo we use a file. For more serious stuff we use postgresql.
More ›