4 ms·
Is task specific compression a thing in real life practical software engineering? As far as reducing data loads go I only came across the keyword "SQL optimizat
by anyfactor 5y ago
Is task specific compression a thing in real life practical software engineering? As far as reducing data loads go I only came across the keyword "SQL optimization".
- hamandcheese 5y agoOne example that comes to mind is some work nvidia has done compressing “video” streams. They do so by capturing your facial movements and reconstructing them on the other side, resulting and massively less bandwidth. https://developer.nvidia.com/ai-video-compression https://developer.nvidia.com/ai-video-compression
- detaro 5y agoDepends how specific you want "specific" to be I guess? I.e. we have loads of compression algorithms for different kinds of data. Whereas here we are looking at almost dataset-specific compression (i.e. the only benchmark is how it works for one specific set of data to compress), and there's a sliding scale between the two ends. Similarly, having to contain the decompression code in the measured result size and it being a relevant contribution is something that only applies in some use cases of compression.
- willis936 5y agoScarcity breeds innovation. Why hand tune a compression scheme for a specific dataset when it is less storage efficient than LZ? You opt to not use LZ when all the memory and compute you have can barely run prefix codes. That's why people still write for the Z80: it's a fun toy.
- psidebot 5y agoI have designed task specific data structures/compression oriented around memory efficiency. In my experience this starts to crop up when datasets get big enough to trigger cost for performance sensitivity. This is especially true for SaaS offerings, where e.g. an ability to stay under the next RAM doubling can result in serious hosting savings.
- morcheeba 5y agoI'm echoing a couple of replies before me, but I'll give concrete examples - MP3, JPEG, and H.264 are all lossy task-specific compressions. Lossless compression includes FLAC and TIFF. For genetic data, HapZipper beats general-purpose compression. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3488212/ https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3488212/ So, yes, actively researched, but you've got to pick a specific task that makes sense. Even small niches are viable; I made a task-specific compressor to strip the essential numbers out of a remote sensor report to make it small enough to squirt to a satellite.
- forgotusername6 5y agoVideo compression algorithms are fascinating. Choosing the correct color space and reducing bit count because the human eye doesn't see color as well, using discrete cosine transform to "group together" the important parts of an image, using diffs from previous and future frames, using diffs from movement in the image. There are so many techniques.
- noduerme 5y agoIt used to be just best practice to shrink down anything over the wire as much as possible. When I started building websites around 1996, every byte counted and you would "optimize" every GIF image carefully to the smallest size you could, taking it down to 256 or 16 or even custom colors - like 6 colors in the VGA spectrum that looked good enough with dithering. It kinda didn't matter from 2010 on. But one area I've written my own specific "compression" methods in, for the last few years, has been in shipping data in and out of webworkers (in-browser or in Node). This is where there's still enough of a performance penalty on a lot of devices for sending 1MB that in use cases where you're spawning lots of workers to run long tasks, it makes sense to trade time to compression for a smaller transfer size.
- voussoir 5y agoHere's another example from a constrained environment: in Andy Gavin's interview with Ars Technica, he talks about a domain-specific compressor they made for storing Crash Bandicoot's animations, because storing each coordinate of each vertex for each frame would have been too large. https://www.youtube.com/watch?v=izxXGuVL21o&t=21m12s https://www.youtube.com/watch?v=izxXGuVL21o&t=21m12s
- justinlloyd 5y agoYep, in video games, especially on older systems, we did an awful lot of that. Storing vector deltas for bone animations rather than the full vector3. Compiled graphics on old PC games I worked on, which both helped in size, but also in speed of blitting to the display. Vector3 are made up of three 32-bit floats, but frequently stored as four 32-bit floats so that everything aligns on word boundaries, but you don't need to do that for data you aren't currently using, so you can save 25% of memory right there. Also, interleaving vector3's for packing and alignment on 3d models. Lots of bitmask manipulations and interleaved SIN/COS and DIV/MUL look-up tables. These last couple of techniques go all the way back to the original Atari Asteroids game from the arcades in the 1970's.
- justinlloyd 5y agoAnything that has limited memory or limited bandwidth, or where the data is much larger than the available memory or bandwidth, e.g. neural compute models at the edge in a IoT device, word databases on limited memory systems, and foveated compression in VR/AR on super high bandwidth connections that still cannot keep up with the 16K video stream given modern video protocols.
- londons_explore 5y agoAnytime you're trying to squeeze more performance out of old/underpowered embedded hardware you'll come across stuff like this... Eg. You work for a doorbell company and the boss says "yo, can we make our doorbell have 6 tunes instead of one, because our competitors are doing that. No, we don't want to change microcontroller".
- ant6n 5y agoAt transit app, we built a domain specific compression for transit schedules. So instead of a city like New York taking 100mb, it takes like 5mb or so. This was a couple years back when data was still more expensive, so one of the things it allowed is simply always downloading schedules for offline availability, instead of having to ask the user when and what to download. Here's a write up (sorry for the cheery tone) with some details https://blog.transitapp.com/how-we-shrank-our-trip-planner-till-it-didnt-need-data-84984ca56663/ https://blog.transitapp.com/how-we-shrank-our-trip-planner-t...
- dahart 5y agoYes! Not only is it a real thing, but with Moore’s Law slowing down, data/compute appetites going up, and the gap between processing speed and memory speed still large and growing, the need for task specific compression is currently going up. Working on GPUs, I see many, and work on some task specific compression ideas as part of my job. The compiler has it’s own ways of compressing code & debug info. The hardware has it’s own ways of compressing textures. A recent feature we built on my team is a compressed encoding for adaptively subdividing curves. All of these things have the primary goal of reducing memory bandwidth, which in turn increases the speed of computation because memory is so frequently the main bottleneck.
- kwhitefoot 5y agoOf course. For instance when you have an embedded controller with limited ROM size and you run out of space, compressing the message strings that are sent to a till roll printer or display might be the only way you can get enough space to add a feature. I had to do this in the early '80s. The alternative was scrapping the boards and redesigning them to allow double the EPROM size but that would have been a lot more costly than writing the decompression routine and manually compressing the strings. It would also have delayed delivery.