Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
faltet
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
Querying 100T rows in a human time scale
(blosc.org)
2 points
by
faltet
4y ago
|
1 comments
2.
▲
by
faltet
4y ago
This is what happens when two nice libraries like HDF5 and Blosc2 cooperate in making a third one (PyTables) way more efficient.
3.
▲
by
faltet
12y ago
Here you have a benchmark based on the MovieLens database: http://nbviewer.ipython.org/github/Blosc/movielens-bench/blo... The results are explained here: http://www.blosc.org/docs/bcolz-
4.
▲
by
faltet
14y ago
I'm not sure why, but yes, I'm seeing these speedups. You can see them too in: http://blosc.pytables.org/trac/wiki/SyntheticBenchmarks and paying attention to compression ratios of 1 (compression disabled). If you find any good reason o
5.
▲
by
faltet
14y ago
Blaze is intended to be a much more general solution than a pure key-value store (although it will be able to tackle this use case too). But yes, the idea under BLZ is close to your projects, namely, leveraging the available resources in y
6.
▲
by
faltet
14y ago
That's right. But even in this case, Blosc, the internal compressor used in Blaze, can detect whether the data is compressible or not pretty early in the compression pipeline, and decide to stop compressing and start just copying (how earl
7.
▲
by
faltet
14y ago
No doubt that much longer :) But the important points to take away are: 1) You can store more compressed data by using the same storage capacities. 2) If data can be compressed, the I/O effort will be less 3) If the compressor is fast enou
8.
▲
by
faltet
14y ago
HDF5 is a very nice format indeed, and in fact, BLZ is borrowing a lot of good ideas from it. However, HDF5 has its own drawbacks, like not being able to compress variable length datasets, the lack of a query/computational kernel or its fl