Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
AdamProut
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
AdamProut
4y ago
Clickhouse is great a single table analytical queries (Group-by + aggregates + filters over a single table). It can match any top tier columnstore database at this, maybe even better at smaller scales, and if that is all you need it will d
32.
▲
by
AdamProut
4y ago
I'm not sure why this is being downvoted? Dividing up memory statically at startup/compile time has some nice benefits described in the article, but it also makes the database less flexible to changing workload requirements. Fo
33.
▲
by
AdamProut
4y ago
I was keeping a tally of how many companies were offering "ClickHouse as a service" at one point last year. I think I got up to 7 or 8. It will be interesting to watch this unfold from a code licensing perspective. Will Clickhou
34.
▲
by
AdamProut
4y ago
This is a reasonable benchmark from the Clickhouse folks for single table anlaytical query performance over smallish data sets (10s of GB of data). Most of the DW vendors are on there. https://benchmark.clickhouse.com/ Tim
35.
▲
by
AdamProut
4y ago
We do still use skiplists for in-memory rowstore indexes. They're great for very high throughput writes (less good for scans). In the years since this blog post most of our engineering effort went towards building a columnstore that
36.
▲
by
AdamProut
4y ago
Yep, lots of real world workloads shard very well. That wasn't my point. My point was that "most" distributed SQL Databases that do sharding will do very well on TPC-C. You can google search for equivalently impressive r
37.
▲
by
AdamProut
4y ago
That's true, I should have been a bit more specific. Each individual query runs within a single shard as far I recall (its been a few years...). There maybe transactions (like new order) that do multiple queries such that the transact
38.
▲
by
AdamProut
4y ago
keep in mind TPC-C scales out in a trivial manner. It was designed before distributed databases existed. The workload itself is perfectly sharded on warehouse_id. All joins happen on this column and the query load is uniformly distributed
39.
▲
by
AdamProut
4y ago
We regularly benchmark the "big 3" Cloud Data warehouses - Redshift, Snowflake and Big Query at SingleStore. Their performance is very close to the same (within 10-20%) on most benchmarks on reasonable sized data sets (10s of TB
40.
▲
by
AdamProut
4y ago
To be clear if fsync() on linux had well defined behavior on errors (no matter the file system used) I wouldn't suggest failing over as a reasonable solution for a distributed database. Its mentioned in the parent video, but fsyncgate
41.
▲
by
AdamProut
4y ago
You could replace step 3) with "pull power plug from host" for the same effect. Think of all the extra disk IO the worlds databases are doing to defend against step 3) :)
42.
▲
by
AdamProut
4y ago
An fsync() failing doesn't necessarily mean there is a disk corruption. I agree all logging and recovery protocols should have different handling for a corruption vs a torn tail of the log for example, but I view that as mostly orthogo
43.
▲
by
AdamProut
4y ago
A nice outcome here for distributed SQL databases that store multiple copies of the data by default is that they can just failover over on an fsync failure and not try to "handle it". If fsync starts failing with EIO there is a
44.
▲
by
AdamProut
4y ago
I talked a bit about some of the advantages in this pretty old blog post: https://www.singlestore.com/blog/what-is-skiplist-why-skipli... Simplicity is really the biggest advantage. Being simple, its much easier to im
45.
▲
by
AdamProut
4y ago
His research on columnstores is seminal. The list of systems above use columnstore storage.
46.
▲
by
AdamProut
4y ago
Couldn't a $ 3 billion valuation cut both ways as far as recruiting is concerned? If the companies current ARR is extremely disconnected from the valuation (say the valuation is 100x ARR), then I would think most experience prospective
47.
▲
by
AdamProut
4y ago
- Firebolt (Hard fork of clickhouse) - Altinity - Gigapipe - Hydrolix - Bytehouse.cloud - https://clickhouse.com/ ("coming soon") - TiDB (Their columnstore is a fork of clickhouse) I stopped trac
48.
▲
by
AdamProut
4y ago
Any idea why Druid performed so poorly though? 100x slower seems odd. I though druid was reasonably good at single table analytics like in this benchmark. Is it the small data size?
49.
▲
by
AdamProut
4y ago
How many Clickhouse as a service offerings exist now? I stopped counting at 7 a few months ago (double.cloud was not on my list).
50.
▲
by
AdamProut
4y ago
I don't know where you draw the line between SQL analytics and SQL data warehousing. I think your typical analytical workload definitely involves more data then this benchmark though. Something like DuckDB is more ideal for this small
51.
▲
by
AdamProut
4y ago
I think your missing my point. The page is entitled "a Benchmark For Analytical DBMS" not "A Benchmark for Single Table Query Execution". Most analytical workloads are more complex then single table queries. I didn'
52.
▲
by
AdamProut
4y ago
It looks like the queries are all single table queries with group-bys and aggregates over a reasonably small data set (10s of GB)? I'm sure some real workloads look like this, but I don't think it's a very good test case to s
53.
▲
by
AdamProut
4y ago
SingleStoreDB is heavily used for this type of app. We used to call this use case real-time analytics (though it has many other names today) [1] https://www.singlestore.com/blog/the-technical-capabilities-... (Disclos
54.
▲
by
AdamProut
4y ago
Clickhouse and Druid are not very good at complex OLAP queries. Clickhouse is pretty upfront about needing to denormalize your schema To avoid distributed joins. Neither are anywhere close to the performance of top DWs on analytical bench
55.
▲
by
AdamProut
4y ago
Filters like this one: WHERE users.name = 'Joe' AND todos.completed = true Have most of the problems I mentioned above unless "many" rows match the filter. Lack of ability to seek to specific rows in the data h
56.
▲
by
AdamProut
4y ago
Yeah, TPC-E is a better (more advanced) OLTP benchmark. TPC-C is trivially scalable by sharding on warehouse id everywhere. The problem with TPC-E is not many companies publish (official or unofficial) results for it, so its not as useful
57.
▲
by
AdamProut
4y ago
For scan heavy apps columnstores will be much much faster. For simple CRUD apps that just want to read/write a few rows at a time (with high concurrency) they have a number of inefficiencies vs a rowstore. - Columnstore often don&
58.
▲
by
AdamProut
5y ago
I suppose. Some problems with mmap() are a bit hard to fix from user land though. You will hit contention on locks inside the kernel (mmap_sem) if the database does concurrent high throughput mmap()/unmap(). I don't follow linu
59.
▲
by
AdamProut
5y ago
I would say that TPC-DS and TPC-H are really table stakes benchmarks for data warehouses at this point in time (maybe they weren't 10 years ago). How to build a database that does well on them is well documented in the literature now[
60.
▲
by
AdamProut
5y ago
> Reductions in generality are a major mechanism for reducing engineering cost in database engines Yeah, just look at snowflake for a great example of this. They outsource the entire bottom half of the database (making data durable) to
More ›