3 ms·
Co-author here -- thanks for linking to our paper. The basic theme is that (a) purely functional implementations of things like video codecs can allow finer-gra
by keithwinstein 9y ago
Co-author here -- thanks for linking to our paper. The basic theme is that (a) purely functional implementations of things like video codecs can allow finer-granularity parallelism than previously realized [e.g. smaller than the interval between key frames], and (b) lambda can be used to invoke thousands of threads very quickly, running an arbitrary statically linked Linux binary, to do big jobs interactively. (Subsequent work like PyWren, by my old MIT roommate, have now done (b) more generally than us.)
My student's talk, which is worth watching and has a cool demo of using the same system for bulk facial recognition, is here: https://www.usenix.org/conference/nsdi17/technical-sessions/presentation/fouladi https://www.usenix.org/conference/nsdi17/technical-sessions/...
Daniel Horn, Ken Elkabany, Chris Lesniewski, and I used the same idea of finer-granularity parallelism at inconvenient boundaries, and initially some of the same code (a purely functional implementation of the VP8 entropy codec, and a purely functional implementation of the JPEG DC-predicted Huffman coder that can resume encoding from midstream) in our Lepton project at Dropbox to compress 200+ PB of their JPEG files, splitting each transcoded JPEG at an arbitrary byte boundary to match the filesystem blocks: https://www.usenix.org/conference/nsdi17/technical-sessions/presentation/horn https://www.usenix.org/conference/nsdi17/technical-sessions/...
My students and I also have an upcoming paper at NSDI 2018 that uses the same purely functional codec to do better videoconferencing (compared with Facetime, Hangouts, Skype, and WebRTC with and without VP9-SVC). The idea is that having a purely functional codec lets you explore execution paths without committing to them, so the videoconferencing system can try different quantizers for each frame and look at the resulting compressed sizes (or no quantizer, just skipping the frame even after encoding) until it finds one that matches its estimate of the network's capacity at that moment.
One conclusion is that video codecs (and JPEG codecs...) should support a save/restore interface to give a succinct summary of their internal state that can be thawed out elsewhere to let you resume from that point. We've shown now that if you have that (at least for VP8, but probably also for VP9/H.264/265 as these are very similar) you can do a lot of cool tricks, even trouncing H.264/265-based videoconferencing apps on quality and delay. (Code here: https://github.com/excamera/alfalfa https://github.com/excamera/alfalfa). Mosh (mobile shell) is really about the same thing and is sort of videoconferencing for terminals -- if you have a purely functional ANSI terminal emulator with a save/restore interface that expresses its state succinctly (relative to an arbitrary prior state), you can do the same tricks of syncing the server-to-client state at whatever interval you want and recovering efficiently from dropped updates.
For lambda, it's fun to imagine that every compute-intensive job you might run (video filters or search, machine learning, data visualization, ray tracing) could show the user a button that says, "Do it locally [1 hour]" and a button next to it that says, "Do it in 10,000 cores on lambda, one second each, and you'll pay 25 cents and it will take one second."
- EgoIncarnate 9y agoThe paper states "ExCamera is free software. The source code and evaluation data are available at https://ex.camera https://ex.camera .", but there doesn't seem to be any source code at the linked site. The github has some source code, but it appears to be incomplete with regards to the binaries also present there. Am I missing something? Did the intentions on this change?
- keithwinstein 9y agoEverything is free software and available at https://github.com/excamera https://github.com/excamera . Not sure what you mean about binaries -- afaik we only have source code checked into our repos. (The video codec is https://github.com/excamera/alfalfa https://github.com/excamera/alfalfa, and the lambda framework is https://github.com/excamera/mu https://github.com/excamera/mu .)
- EgoIncarnate 9y agoI'm referring to https://github.com/excamera/excamera-static-bins https://github.com/excamera/excamera-static-bins . It's not clear where the build scripts, etc are for these projects.
- TD-Linux 9y agoMany encoders have a "two pass" mode that saves statistics from a first, fast pass to guide the second pass. Usually these are very simple statistics, used to choose a quantizer and frame type in the second pass. Your method feels like a rather extreme version of this, that can guide the whole RDO search with a very large state. And, of course, rarely are first passes done in parallel. It's super exciting to see people looking at this - I think there's a lot of untapped potential. The WebRTC case looks interesting too, though there you need extremely low latencies (1 frame) so I'm curious to find out how you tackle that. One downside with your approach is that you have to code many keyframes that are eventually thrown away. Keyframes are often expensive to encode because their bitrate is so much higher. Have you considered "synthesizing" fake keyframes somehow, such as doing an especially fast and stupid encode for them? Also, artifacts from keyframes tend to be a lot different from inter predicted frames, so would encoding a keyframe, plus a second frame, then throwing away both yield a quality improvement at the final pass (at the cost of more wasted computation)?