3 ms·
I think the MPEG people and the authors of pl_mpeg probably just differ on how "very hairy" it is to represent the byte-by-byte state of a single-threaded decod
by keithwinstein 7y ago
I think the MPEG people and the authors of pl_mpeg probably just differ on how "very hairy" it is to represent the byte-by-byte state of a single-threaded decoder. In other words, how hard is it to create a continuation object that lets the decoder run out of buffer at any byte location, return to the caller, and later resume from the same place when more bytes are available? The MPEG people were mostly thinking about hardware decoders, but this is not that hard in software -- and it's been done successfully by every major decoder implementation afaik.
The page writes: "if we're in the middle of decoding a video frame and the buffer doesn't have any more bytes yet (e.g. because it's streaming from the net) we would need to pause the decoder, save its exact state and later, when enough bytes are available, resume it again. Of course this isn't particularly difficult to achieve using threads, but if we want to stay single threaded it gets very hairy."
But I'm pretty sure the MPEG people would just say, "look, this is not that hard and was a solved problem 20 years ago. We're not going to introduce a 1-frame delay at the encoder (which is what you want us to do by putting a frame length in the PES header) but we're also not making you introduce a 1-frame delay at the decoder. Just do what libmpeg2 has been doing since 1999: make a state object that represents the byte-by-byte evolving state of the decoder. Update the state object as you go. If you run out of bytes in the buffer, return STATE_BUFFER to the caller. When the caller gives you more bytes later, use the state object to resume from where you left off."
"Here's what that object looks like for MPEG-2: https://github.com/cisco-open-source/libmpeg2/blob/master/libmpeg2/mpeg2_internal.h#L67-L155 https://github.com/cisco-open-source/libmpeg2/blob/master/li...
Yes, it's not beautiful, and yes, it would be more elegant if these variables were broken out into different scopes (current block, current slice, current picture, current sequence), but this is basically what you're in for and it's not that hard. And it avoids a 1-frame delay at either end. (And no, you don't need a 1-frame delay just because you're demuxing from a transport stream either -- just throw each TS packet into the video ES decoder as soon as you get it.)"
- phoboslab 7y agoThanks for the explanation! It hadn't occurred to me that an encoder would spit out packets for part of a frame, while it's still encoding. For what it's worth, ffmpeg's public API - at least the one I used in a previous project[1] - only produces full frames. The MPEG-2 state object you linked to looks a lot like the private data of my decoder already[2]. I wonder if there's any restriction on when a packet may be concluded. I.e. do MPEG-PS packets have to contain full slices, or can they be cut off in the middle of slice? The "hairy" part with my current design would be to reproduce the call stack. Again, if the decoder would live in its own thread, it would be a no-brainer. > and it's been done successfully by every major decoder implementation afaik As far as I can tell, ffmpeg's decoder does not allow for this. It always searches for the next picture's START_CODE before starting to decode the frame. Similarly, the libmpeg2 source you linked to doesn't seem to provide any functionality to resume decoding from anywhere in the stream either!? The NEEDBITS and DUMPBITS macros just assume there's always more data. [1] https://github.com/phoboslab/jsmpeg-vnc/blob/master/source/encoder.c#L76 https://github.com/phoboslab/jsmpeg-vnc/blob/master/source/e... [2] https://github.com/phoboslab/pl_mpeg/blob/master/pl_mpeg.h#L1775-L1825 https://github.com/phoboslab/pl_mpeg/blob/master/pl_mpeg.h#L...
- keithwinstein 7y agoWell.... now that I've gone to read the code again, you're absolutely right that I was too hasty in calling it the "byte-by-byte" state. It's more like the "slice-by-slice" state. libmpeg2 goes header-by-header, so, it can process each individual slice without waiting for the next picture start code, but it does buffer up a whole slice before starting any work. If you just give it a single byte (or any number of bytes that doesn't include some subsequent start code, including a sequence_end_code for the end of the whole video), it just copies it to an internal buffer and then asks for more until it sees the beginning of the next slice or some other header. That's why NEEDBITS and DUMPBITS don't have to bail out in the middle -- by the time you get there, you know they have a whole slice to play with. So, yes, libmpeg2 does go start-code (or sequence_end_code) by start-code -- but not a picture start code. ffmpeg/libavcodec is a wrapper around like 75+ different decoders, so I'm not too surprised if they have to go with a least-common-denominator interface. In general an MPEG-2 TS or PS packet is just a fixed size packet and doesn't have to be aligned with any ES syntax element. Typically the PES packets (the much larger packets encapsulated in PS/TS packets) do contain exactly one video picture (i.e. the data_alignment_indicator is set on every video PES packet), but even this isn't formally required. Note that the PES packet header also includes an optional length field that would do what you want (but it's optional, in part to accommodate encoders that don't want to buffer the whole image before starting to encode pixels). You might be interested in our TS/PES demuxing code that wraps libmpeg2/liba52 and tries to maintain a/v sync in the presence of arbitrary corruption -- it's more than half the size of your entire decoder! https://github.com/StanfordSNR/puffer/blob/master/src/atsc/decoder.cc https://github.com/StanfordSNR/puffer/blob/master/src/atsc/d...
- kevin_thibedeau 7y agoReplace the call stack with a state machine. Then you don't have to monkey around with coroutines to get back to where you left off.