Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gianm
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Building a sentiment analysis application with ChatGPT and Apache Druid
(imply.io)
1 points
by
gianm
4y ago
|
0 comments
2.
▲
by
gianm
4y ago
It does seem odd, especially since in real world cases I'm more accustomed to seeing Druid and ClickHouse be in the same ballpark of performance. Sometimes one is somewhat faster than the other. But in my experience that's more li
3.
▲
by
gianm
4y ago
This is impressive work: it's time consuming to set up and benchmark so many different systems! Impressiveness of the effort notwithstanding, I also want to encourage people to do their own research. As a database author myself (I work
4.
▲
by
gianm
6y ago
I'm a committer on Apache Druid and generally a big fan of observability. I'm glad that you found Druid useful in building this! A tip, if you aren't already doing it: with metric and trace data, it helps a ton to set up part
5.
▲
by
gianm
7y ago
Those are good questions. IMO Druid is most well-differentiated if you want to power an online, real-time, high-concurrency analytical application at scale. It is the use case Druid was originally designed for and still the one where the pr
6.
▲
by
gianm
7y ago
Hey Mani. Druid committer here. It actually is a column store! The project makes a big deal about its ability to do indexes and pre-aggregation because those are important capabilities and, while not unique, are also not universally support
7.
▲
by
gianm
7y ago
Logically, an array of booleans and a set of integers are equivalent. So in the Druid developer community we usually use the terms interchangeably. But to be precise, our indexes are all stored as bitmaps and compressed with bitmap compress
8.
▲
by
gianm
7y ago
Druid committer here. (Also, I think we've met before in SF!) One thing I wanted to add with regard to performance. Druid does indeed get a big boost from the fact that it uses inverted indexes for filtering. It also gets a boost from
9.
▲
by
gianm
7y ago
Druid committer here. Fwiw, Druid was designed to run on huge clusters and that really shows up in the multi-process architecture. The idea is that if you separate the components needed for ingestion, historical processing, query routing, a
10.
▲
by
gianm
8y ago
Check out Druid [1], an open-source analytical database with tightly-coupled storage and processing engines designed for OLAP. In particular it implements a memory-mappable storage format, indexes, compression, late tuple materialization, a
11.
▲
by
gianm
8y ago
Fwiw, more recent versions of Druid have a no-rollup mode that does ingestion row-for-row. It ended up being useful for cases where you _do_ care about every row, maybe because you want to retrieve individual rows or maybe because you don&#
12.
▲
by
gianm
10y ago
Druid committer here, happy to answer any questions!
13.
▲
Imply raises seed round from Khosla Ventures for Druid
(venturebeat.com)
3 points
by
gianm
11y ago
|
0 comments
14.
▲
by
gianm
12y ago
The post says Pulsar can use Druid as a metrics store, so that workload should be doable. Druid is meant for exactly that sort of thing (fast aggregates with ad-hoc filters).