Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jkbonfield
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
jkbonfield
4y ago
Well yes it was one file, but it was stated as being good on text and enwik8 is a pretty standard test corpus for text compressors. I could have done more, but it somewhat vindicated what I was saying really. It has a very similar core to
2.
▲
by
jkbonfield
4y ago
It doesn't compare itself against bsc, which feels a bit poor IMO given it's using Grebnov's libsais and LZP algorithm (he's the author of libbsc). On my own benchmarks, it's basically comparable size (about 0.1% sm
3.
▲
by
jkbonfield
5y ago
As the author of the CRAM implementationn of rANS, I can say that these sort of articles aren't helpful. Clearly my work predates this by several years, so there is nothing here which can realistically impact on CRAM, however fear alo
4.
▲
by
jkbonfield
7y ago
I know I'm not the best at anything I do, but when combined the overlap makes a niche for me that has so far worked out well. Pure luck frankly. However that is only because the skills I'm thinking of aren't in themselves
5.
▲
by
jkbonfield
8y ago
Some of the MPEG-G authors are experts in genomics data compression, while others are experts in video compression. It should, in theory, be a good mix. MPEG are also well aware of the prior art. The authors of various existing state of
6.
▲
by
jkbonfield
8y ago
Firstly, GA4GH has commercial members as well as academics, and all collaborate together to produce file formats, standards and protocols. Secondly you missed out a key part of funding - precompetitive alliances. Eg see the Pistoia Allianc
7.
▲
by
jkbonfield
8y ago
Show me where they notified others taking part of their patents or intent to patent. They sought out academics and invited them to take part. Yes I was naive, but I also felt rather mislead. The GenomSys patents aren't even listed in
8.
▲
by
jkbonfield
8y ago
I am the author of an implementation, although not the author of the file format itself. Although yes that it is still a fair point if you look at just the one blog post. However there are a series of them where I clearly explain the pr
9.
▲
by
jkbonfield
8y ago
Disclaimer - I am the author of the blog. There are no "comparisons with" CRAM in the MPEG-G preprint, only comparison between CRAM and DeeZ, taken from the DeeZ paper. Those comparisons are fair and correct, but obviously were d
10.
▲
by
jkbonfield
8y ago
It's hard with huffman given you need to deal with N. Realistically you'll end up with 3 bases at 2 bits, 1 and 3 and the other 3 being a prefix for everything else (N, ambiguity codes, etc), so somewhere averaging close to 2.3 b
11.
▲
by
jkbonfield
8y ago
To be "fair", their latest revision of the paper includes some figures from a 4 year out of date version of CRAM, which is an improvement on the 10 year old outdated format they used initially. ;-) Disclaimer, that's my blog