Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
daniel_rh
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
31.
▲
by
daniel_rh
8y ago
I'm a big corrode fan and I used it (plus a really simple python script https://github.com/dropbox/rust-brotli/blob/master/uncorrode... ) to help translate the brotli encoder https://gith
32.
▲
by
daniel_rh
8y ago
DivANS author here: the compression benchmarks all measure seekability to the nearest 4MB chunk, but the current lib itself doesn't support seekability out of the box, yet. However, it would be trivial to reset the encoder at each 4mb
33.
▲
by
daniel_rh
8y ago
It has some similarities to zpaq, especially with the -mixing=2 parameter. With -mixing=2, DivANS runs 2 models to estimate the upcoming nibble or byte and dynamically chooses the best model based on past performance. One of the two models
34.
▲
by
daniel_rh
8y ago
To compress files for a system like Dropbox we would recommend a) check if the file can compress with https://github.com/dropbox/lepton . Lepton has some new flags that can help it compress a wider range of files, so i
35.
▲
by
daniel_rh
8y ago
The crate is built with cargo build --release --features=simd --target=wasm32-unknown-unknown which aliases DefaultCDF16 type to the SIMD types. However since WASM doesn't support vectorized instructions yet, I believe LLVM translates
36.
▲
Show HN: DivANS – Rust simd compression algorithm demo in WASM in the browser
(dropbox.github.io)
36 points
by
daniel_rh
8y ago
|
5 comments
37.
▲
by
daniel_rh
8y ago
Blog co-author here: With the settings in the blog post, the DivANS algorithm skews towards saving space over speed. Brotli skews heavily towards decompression speed without sacrificing much ratio. The reason we focus a lot on comparing to
38.
▲
by
daniel_rh
9y ago
I'm the author of https://github.com/dropbox/rust-brotli and I certainly ran into the issues mentioned in the article, and the initial version of my brotli decoder and encoder were each almost 10x slower. I also w
39.
▲
by
daniel_rh
9y ago
The researchers specifically tested phone face up vs phone face down and noticed no statistical difference, I believe.
40.
▲
by
daniel_rh
9y ago
I thought you could use memoryview over a string to get rid of that allocation even in 2.7 https://docs.python.org/3/library/stdtypes.html#memoryview
41.
▲
by
daniel_rh
9y ago
GPL3 means it's not very reusable in any projects that aren't GPL3. Have you considered a BSD or MIT license? or CC0 if you want to get code reuse? Or at the very least LGPL2
42.
▲
by
daniel_rh
10y ago
SIMD support in Rust is still very early and required unsafe mode for SSE, and lepton makes heavy use of SSE intrinsics. Now I do see http://huonw.github.io/simd/simd/ was being developed in August 2015, but it se
43.
▲
by
daniel_rh
10y ago
rANS is really cool, but I think we tried it too early in the lepton research phase while some of the other ideas were still brewing... We might revisit rANS now that v1.0 is out!
44.
▲
by
daniel_rh
10y ago
Hi GrantS: That's pretty close to how it works. We always use the new prediction, even where it's worse than JPEG though (very rarely), to stay consistent. As for having the mega-model that predicts all images better: well it turn
45.
▲
by
daniel_rh
10y ago
It's very easy to make a JPEG file that will not compress at all. Luckily it might look like snow from a television set rather than a typical image produced by a camera. We live in a world where it is common for blue sky to occupy a po
46.
▲
by
daniel_rh
10y ago
Thanks for testing this! For archival purposes, be sure to check the exit code of the lepton binary after compressing each JPEG: the default parameters only support a subset of JPEGs (you need to pass -allowprogressive and -memory=2048M -th
47.
▲
by
daniel_rh
10y ago
Do you have a link to the code? Is it open source? Sounds pretty cool :-)
48.
▲
by
daniel_rh
10y ago
That's the idea behind the algorithm, yes. And since it's lossless, every original bit is preserved. The same idea could be applied on the Desktop client instead of on the server, which would save 22% of the bandwidth as well and
49.
▲
by
daniel_rh
10y ago
Your argument is certainly reasonable. One case where it might matter is if SECCOMP weren't available on all platforms that needed to decompress data. Also: a vulnerability could decide to only strike after a certain clock date. In tha
50.
▲
by
daniel_rh
10y ago
The file's sha256sum can be verified before the file is sent to any users, so there's no chance of RCE there, even with a hypothetical C brotli-- but a reproducible decompression is key. Additionally, if you want to do the decompr