6 ms·
Neat. I spent a month trying to use sphere packing approaches for a better compression algorithm (I had a large amount of vectors, they were grouped through clu
by Scene_Cast2 1y ago
Neat. I spent a month trying to use sphere packing approaches for a better compression algorithm (I had a large amount of vectors, they were grouped through clustering). Turned out that theoretical approaches only really work for uniform data and not any sort of real-world data.
EDIT: groped -> grouped
- Gregaros 1y ago_May_ be a case for extending out what has been explored by theory to cover more useful ground (or not, depending on whether real-world usecases like yours are too heterogenous for effective general techniques).
- soulofmischief 1y agoYou really shouldn't grope your vectors.
- dotancohen 1y agoRoger, Rodger. Over, Oveur.
- GuinansEyebrows 1y agoyou're Kareem Abdul-Jabar!
- dotancohen 1y agoI'm sure you've already explored this, but is there some precompression operation that you could do to the vectors such that they're no longer sparse, and therefore relatively uniform?
- Scene_Cast2 1y agoThey weren't sparse, they were dense but the "density" was quite non-uniform (think typical learned ML vectors). Not too far from an N-dimensional gaussian (I ended up reading research on quantizing Gaussian distributions, but that didn't help either as we didn't have a perfectly gaussian thing).
- sdenton4 1y agoVAE objectives are useful for pushing embeddings into a Gaussian distribution. Here's some work on low-latency neural compression that you might find interesting: https://arxiv.org/abs/2107.03312 https://arxiv.org/abs/2107.03312
- dotancohen 1y agoI see, thank you. Another thing that I'm sure you explored, and I'd love to hear how it went, would be to rearrange the elements in the vectors such that perhaps the denser parts could be more contiguous, and the sparser parts could be more contiguous, on average. That sounds like something that would be easier to compress. Were the distributions such that a rearrangement like this might have been possible? Or were they very evenly distributed? I.e. could you have rearranged a Gaussian-like distribution into a Poisson-like distribution?
- Scene_Cast2 1y agoWhat ended up launching is a fancy product quantization based on k-means. Some of the tricks were storing magnitude separately (i.e. removing the mean) and rearranging dimensions based on variance (and/or rotation based on PCA) for PQ to work better. I also remember trying to fit a distribution so that I can generate synthetic data (not for a lack of data, but more for understanding the problem space better). The synthetic data quantized pretty differently - my guess is that it's because of random areas of density and sparsity. I'm not quite following your exact rearranging idea though. Not sure if the above answers the question.
- dotancohen 1y agoI was just wondering, it seems like an interesting problem to explore.
- derf_ 1y agoIf you start from N-dimensional Gaussian-distributed vectors and normalize them, you wind up with uniformly distributed vectors on an (N-1)-dimensional hypersphere. Of course, sphere-packing on the surface of a hyperspehere is its own fun problem, and if your data is not actually Gaussian may still not be exactly what you want, but coding the vector magnitude separately from the vector direction is probably a good start.
- lupire 1y agoIt's usually the case that the low hanging fruit in a decades old commercially valuable field have already been picked.
- hansvm 1y agoThe usual trick is to use domain-specific knowledge to translate that asymmetry to uniformity. E.g.: Suppose the data has high-order structure but is locally uniform (very common, comes about because of noise-inducing processes). Compute and store centroids. Those are more uniform than your underlying data, and since you don't have many it doesn't really matter anyway. Each vector is stored as a centroid index and a vector offset (SoA, not AoS). The indices are compressible with your favorite entropic integer scheme (if you don't need to preserve order you can do better), and the offsets are now approximately uniform by assumption, so you can use your favorite sphere strategy from the literature.
- tsunamifury 1y agoIntelligence is compression