4 ms·
> Apologies if this is a noob question, but could you improve it with training? Considering all other things like GPU, code quality etc. equal, the speed here
by esquire_900 5y ago
> Apologies if this is a noob question, but could you improve it with training?
Considering all other things like GPU, code quality etc. equal, the speed here mainly depends on the size of the network. Larger networks result in better performance (logarithmic), but also slower speeds (linear). There are some tricks like pruning to create significantly smaller networks with similar-ish performance, but good models are often large.
Unrelated to your question, what nobody mentions is that the decompression side also needs to complete network. With 57M and 187M parameters, those files are going to be quite large, > 100MB for sure. That completely annihilates the performance wins for 1 time transfers.
- 317070 5y agoNo, as I understand it, the network does not need to be transmitted. The decoder learns on the same stream as the encoder, so you do not need to transmit network parameters.
- ironSkillet 5y agoIf I compress a large file at location A using this algorithm, and want to send it to location B, how does location B know how to decompress it?
- clavigne 5y agosee the (current) top comment https://news.ycombinator.com/item?id=27244810 https://news.ycombinator.com/item?id=27244810 it's a very clever use of symmetry.
- vidarh 5y agoIt's clever, but pretty much "standard" in compression. Earliest I'm aware of that used symmetry in the encoder and decoder to prevent explicitly transferring the parameters this way was LZ78 (Lempel, Ziv; 1978), but there could well be predecessors I'm not aware of. LZ78 used it "just" to build a dictionary, but the general idea of using symmetry is there. (A fascinating variant on this is Semantic Dictionary Encoding, by Dr Michael Franz in his doctoral dissertation (1994), that used this symmetry to compress partial ASTs to cache partially generated code for runtime code generation)
- e12e 5y agoSee https://news.ycombinator.com/item?id=27244810 https://news.ycombinator.com/item?id=27244810 As I understand it, when starting up (beginning of stream) - the network does not compress, then guesses the next byte at every turn determininstically from input so far. Decoding can "read" the fist byte, then proceed to learn and guess the second byte and so on.
- Yoric 5y agoBasically the same way your favorite codec knows how to decompress. It builds the predictor (in this case, a NN) from the known input. The predictor then assigns a symbol (or a series of symbols) to the next few bits - higher probabilities need fewer bits (I haven't checked for this specific technique, but there are techniques that don't even require an integer number of bits, e.g. range encoding).