29 ms·
H.266/Versatile Video Coding (VVC)
- KitDuncan 6y agoCan we all just agree on using AV1 instead of another patent encumbered format?
- hirako2000 6y agoHow does x264 etc get away with it?
- clouddrover 6y agoThey don't get away with it. The are two separate issues: the software copyright license and the H.264 patent license. x264 itself is licensed under the GPL: http://www.videolan.org/developers/x264.html http://www.videolan.org/developers/x264.html But if you use an H.264 encoder or decoder in a country that recognizes software patents then you need to buy a patent license if your usage comes under the terms of the license: https://www.mpegla.com/programs/avc-h-264/ https://www.mpegla.com/programs/avc-h-264/
- p_l 6y agoThey are compliant with MPEG patent licenses - they distribute source code you have to compile yourself and for non-commercial use. If you build x264 or other open source implementation of h.264/h.265, and embed it for example in commercial video conferencing software/appliance, you have to pay patent licensing fees for that product. It's also why Firefox downloads a blob from Cisco to handle MPEG-4 video - Cisco covers the licensing for distribution et al.
- jiggunjer 6y agoHow much does Cisco pay (to MPEG?) for that?
- detaro 6y agoIf I understand correctly, around 10 million USD per year (according to https://www.mpegla.com/wp-content/uploads/avcweb.pdf https://www.mpegla.com/wp-content/uploads/avcweb.pdf that's the cap, and from what I've heard Cisco is selling enough actual products to hit that cap, so providing their software for free to everyone else doesn't cost them any extra in licensing fees, just hosting and such)
- mekster 6y agoWhy is Cisco being nice about it? (Besides why not?)
- toast0 6y agoI don't recall the specific timing of the release, so this might not line up, and I have no inside knowledge, just public information. Cisco has some products which use compressed video in a browser setting. It would be useful if all browsers supported a good codec. Individually downloaded codec plugins suck, because installing is iffy. Therefore, give something away which doesn't cost licensing money to make your existing licensed products more usable. And get some good feels on the interwebs.
- eqvinox 6y agoBecause it forces their competitors (in the video conferencing business) to take similar costs.
- brtech99 6y agoBecause several years ago, there was a fight over mandatory to implement video codecs in WebRTC. It was VP8 vs H.264. The biggest thing VP8 had going for it was no royalty payments. Cisco wanted H.264 because all of their devices supported H.264 and none supported VP8 and they already paid the royalties. So Jonathan Rosenberg, then CTO of the division of Cisco that managed this part of the business arranged to have Cisco cover the royalty payments for anyone implementing the WebRTC standards. That wasn't enough, and WebRTC requires both VP8 and H.264 as MTI codecs.
- moonchild 6y agoCould firefox download x264 and compile it on-demand?
- ksec 6y agox264 is for encoding only.
- realusername 6y agoSoftware patents are not valid everywhere, x264 is brought up by VideoLAN in France, where software patents don't apply like in the rest of the EU.
- jacobush 6y agoThey sort of half not apply in Sweden too.
- anticensor 6y agoWhat do you mean by half valid?
- jacobush 6y agoSoftware patents are valid only if part of a larger invention, the algorithm itself cannot be patented. However, patent clerks have also from time to time registered algorithmic patents.
- pault 6y agoSo if a company in france writes open source software that infringes on a US software patent, and a company in the US bundles that source code in a product, is the US company liable for damages?
- pas 6y agoYes. The company distributing the work has to make sure it has licenses for any required parts.
- realusername 6y agoYes, that's exactly how it works, the US company is responsible for following the US laws.
- hirako2000 6y agoJust found av1 is about 20 to 30% more efficient than h265. I guess there was no reason to use patented algs, but h265 is now significantly more efficient than av1. I would still take freedom over patented software
- ch_sm 6y ago> Just found av1 is about 20 to 30% more efficient than h265 > […] but h265 is now significantly more efficient than av1. What did you mean?
- smiley1437 6y agoNot OP but I think he meant "h266 is now significantly more efficient than av1"
- sly010 6y agosecond h265 was probably meant to be h266
- deleted 6y ago[deleted]
- zionic 6y agoAV1 beats h.265 by 20-30%, h.266 beats h.265 by 50%. Honestly for that little of an improvement I'll go with AV1.
- freeone3000 6y agoReally? Because by that metric, H.266 is as far ahead of AV1 as AV1 is ahead of h.265.
- randall 6y agoThat's the direction everyone is going, so I think h.266 is an attempt to give vendors pause before moving to av1. AV1 will probably win in most circumstances (big tech) but is unlikely to win where there are big gains to be had by reducing file size (broadcasters with gigantic libraries). Broadcasters are also used to paying a lot and not getting much.
- otterley 6y agoNo, because the market is more than happy to pay a few cents or dollars per device to get better compression and lower transmission bandwidth. This observation has held true consistently in the 3 decades since compressed digital media was invented.
- jimktrains2 6y agoIt's not that the market is happy to pay more, it's that there is essentially no choice.
- inscionent 6y agoA captive market is a happy market, no?
- rbinv 6y agoNot for consumers, obviously?
- otterley 6y agoThis is not a captive market. If someone were to invent a free codec that performed similarly and had both software and hardware reference implementations, the market would adopt it very quickly. This is a market that is voluntarily paying for perceived value.
- jimktrains2 6y agoPlease show me where I can pick between similar consumer devices where there supported codecs are easy to find, or, better yet, I can pay extra for non free codecs.
- otterley 6y agoThe few HNers who actually care about these things do not make a market that vendors think is worth serving, most likely because it would be unprofitable. There’s not going to be a market if sellers don’t find it profitable.
- donatj 6y agoHonestly, the first step in this is getting the ffmpeg av1 library to a good usable place. It's currently so slow as to be near unviable. I'd happily switch when it becomes a usable option.
- zionic 6y agoIsn't that largely dependent on hardware acceleration from CPU manufacturers? Or is ffmpeg always software encoding?
- oefrha 6y agoFFmpeg isn’t “always software encoding”, that statement doesn’t make much sense since FFmpeg/libavcodec is more of an interface and you can add support for any external encoder/decoder, hardware accelerated or not. However, FFmpeg’s builtin encoders and the most popular external encoders including x264 and x265 are all software encoders. There are hardware accelerated encoders from GPU vendors, e.g. the nvenc encoder for H.264 and H.265 is available for use on Nvidia GPUs if your FFmpeg is compiled against CUDA SDK. It’s a lot faster than x264 and x265 on comparable settings but results are usually a bit worse. You were probably thinking about hardware decoding though.
- donatj 6y agoI'm curious why hardware encoding is generally worse? All my experiments with it (h264/h265) have lead to significantly lower quality output to the point that I've avoided it for any final outputs, but I always assumed I was doing something wrong.
- oefrha 6y agoI’ve written various applications on top of FFmpeg but never looked into encoder technicalities, so I don’t know if and why GPU encoding has fundamental limitations. I heard nvenc has vastly improved on RTX cards; I’m still rocking an old GPU so can’t verify that, but presumably that means there’s no fundamental weakness and GPU encoding is just playing catch up?
- MayeulC 6y agoI wonder what would happen if ffmpeg made the choice to not implement a decoder for it. Or maybe just not a encoder? Can we agree not to work on such projects? I feel that the lack of a good open-source encoder/decoder would spell the death of most codecs nowadays. That would also teach Fraunhoffer about it. Of course, everyone is free to scratch their itch. And the bigger the void, the more itchy it gets. Luckily, we still have AV1.
- donatj 6y agoHuh. I wonder how encoding speeds compare. I rarely chose h265 over h264 because similar levels of visual quality took massively more time.
- corty 6y agoMy guess is, encoding speed will be worse. Video codecs for non-realtime applications are optimized for size and acceptably cheap playback. Encoding performance doesn't really matter since you encode only once but play and store more often.
- ctdeneen 6y agoBut for most situation you encode once and play multiple times. Wouldn't it be better to reduce storage and bandwidth costs with a smaller file (assuming the same quality)?
- donatj 6y agoIt's a trade off. When I have a batch of 40 videos and encoding h264 takes 20 minutes per video and h265 takes 4 hours means the difference between 13 days and 160 days. The latter isn't practical, I'll eat the couple hundred MB in order to save a lot of time.
- waterhouse 6y agoI think by "13 days and 160 days" you mean "13 hours and 160 hours".
- cagenut 6y agothat sounds great, but this is a press release with no real technical details. can anyone in the know add some context? for instance, whats the tradeoff? I assume more CPU? webrtc based video chats are all still using h264, did they not adopt 265 yet for technical or licensing reasons? what is the likelihood of broad browser support for h266 anytime soon?
- rjsw 6y ago> webrtc based video chats are all still using h264, did they not adopt 265 yet for technical or licensing reasons? Is that with x265 built into both browsers ? I build it into mine but I don't think it is the default for ffmpeg.
- spuz 6y agoNot sure why this was downvoted. These all seem like very reasonable questions that others here might be able to answer.
- TheRealSteel 6y agoH.265/HEVC takes about ten times as much computation to encode than H.264 [1], so H.264 still has legitimate technical use cases, even with licensing/patents aside. This makes it great for a company like Netflix or YouTube, but less good for one-to-one and/or battery sensitive use cases like video calls. However, specialized chips help, and some mobile devices can record in HEVC in real time (mine from 2019 can). I believe current smartphones have HEVC encoding hardware, but I'm struggling to find a source for that right now. I haven't seen the details of this new codec yet, but it's quite possible it also has a large encoding cost which will make it better suited to particular use cases, as opposed to a blanket upgrade. [1] http://www.praim.com/en/news/advanced-protocols-h264-vs-h265/ http://www.praim.com/en/news/advanced-protocols-h264-vs-h265...
- becauseiam 6y agoiPhone 7 onwards[1], Qualcomm Snapdragon 610 onwards[2], and Intel Skylake and later CPUs[3] can all encode and decode H.265 in hardware to varying profile levels. 1: https://support.apple.com/en-gb/HT207022 https://support.apple.com/en-gb/HT207022 2: https://www.qualcomm.com/snapdragon/processors/comparison https://www.qualcomm.com/snapdragon/processors/comparison 3: https://trac.ffmpeg.org/wiki/Hardware/QuickSync https://trac.ffmpeg.org/wiki/Hardware/QuickSync
- eigenvalue 6y agoCan anyone verify if this is a real number? It’s possible sometimes to make surprising claims (such as 50% lower size) by relying on unusual or unrealistic situations. I would rather if they use a standard set of test videos with different content and resolutions, and some objective measure of fidelity to the original, when quoting these percentages. But if the 50% number is real, then that is truly remarkable. I wonder how many more CPU instructions are required per second of decoded video compared to HEVC.
- ajross 6y agoIf one trusted numbers like this, and followed a chain of en vogue codecs back through history, you'd expect that a modern codec would produce files sizes like 3% of MPEG-2 on the same input data. It's all spin. I'm sure it does better, but I'm equally sure it'll turn out to be an incremental benefit in practice. > I wonder how many more CPU instructions are required per second of decoded video compared to HEVC. CPU cycles are cheap. The real cost is the addition of yet another ?!@$!!#@ video codec block on every consumer SoC shipped over the coming decade. Opinionated bile: video encoding is a Solved Problem in the modern world, no matter how much the experts want it to be exciting. The low hanging fruit has been picked, and we should just pick something and move on. JPEG-2000 and WebP failed too, but at least there it was only some extra forgotten software. Continuing to bang on the video problem is wasting an absolutely obscene amount of silicon.
- syshum 6y agoThe video codec battle is not about file size, it is about streaming bandwidth.
- yjftsjthsd-h 6y agoAre those different? If I chop up a video file into chunks, I'm streaming it, and if I save a stream I have a file. With a buffer, I would expect the sizes involved to be identical. (Although without a buffer I'd expect streaming to be worse)
- clouddrover 6y ago> A uniform and transparent licensing model based on the FRAND principle (i.e., fair, reasonable, and non-discriminatory) is planned to be established for the use of standard essential patents related to H.266/VVC. Maybe. On the other hand, maybe not. Leoanardo Chiariglione, founder and chairman of MPEG, thinks MPEG has for all practical purposes ceased to be: https://blog.chiariglione.org/a-future-without-mpeg/ https://blog.chiariglione.org/a-future-without-mpeg/ The disorganised and fractured licensing around HEVC contributed to that. And, so far, VVC's licensing looks like it's headed down the same path as HEVC. Maybe AV1's simple, royalty-free licensing will motivate them to get their act together with VVC licensing.
- DrBazza 6y agoNaively hoped I'd read 'this will be released to the community under a GPL license' or similar. Instead found the words 'patent' and 'transparent licensing model'. I appreciate that it costs money and time to develop these algorithms, but when you're backed by multi-billion dollar "partners from industry including Apple, Ericsson, Intel, Huawei, Microsoft, Qualcomm, and Sony" perhaps they could swallow the costs? It is 2020 after all.
- yewenjie 6y agoIs H.265 released under a GPL-like licenses? If not, how do softwares like Handbrake use it?
- speedgoose 6y agoThey use ffmpeg, which is developed by people who do not care about software patents because it doesn't apply to them.
- pault 6y agoWhy don't the patents apply to them?
- speedgoose 6y agoI don't know for all contributors, but the creator is French and there is no software patents in France.
- jcranmer 6y agoAll of the French patents listed in this license pool care to disagree: https://www.mpegla.com/wp-content/uploads/avc-att1.pdf https://www.mpegla.com/wp-content/uploads/avc-att1.pdf
- speedgoose 6y ago
- Unklejoe 6y agoIt's interesting that they are able to continue improving video compression. You'd think that it would have all been figured out by now. Is this continued improvement related to the improvement of technology? Or just coincidental? Like, why couldn't have H.266 been invented 30 years ago? Is it because the computers back in the day wouldn't have been fast enough to realistically use it? Do we have algorithms today that can compress way better but would be too slow to encode/decode?
- aey 6y agoCompression is AI. It’s never going to be “all” figured out.
- 0-_-0 6y agoAnother way if saying it is that compression is understanding.
- jacobush 6y agoLossy compression is, I feel compelled to add.
- aey 6y agoActually both! Arithmetic coding works over any kind of predictor.
- milansuk 6y agoEnd-credits are just text. So it should be possible to put it through OCR and save only text, positions, and fonts. And the text is also possible to compress with a dictionary.
- gsich 6y agoTrue, but end credits take very little space compared to the rest of the movie.
- m3kw9 6y agoIt will be first adopted by pirates for sure
- syshum 6y agoH264 seems to still be the preferred codec in this space, even though H265 is a smaller file size. Largely due to the CPU over head of H265, though I am not sure why more people do not use GPU encoding over CPU Encoding, I have never been able to notice the difference visually
- ctdeneen 6y agoNvidia's NVENC doesn't support CRF, one of the more popular methods of rate control during encoding.
- gsich 6y agoNVENC is too low quality for the scene.
- seanalltogether 6y agoFrom what I can tell from a cursory look at some popular tv show torrents, everything 1080p and below is still h.264, with everything above that running on h.265
- zionic 6y agoNot for 4K rips. I see tons of x265 10 bit encodes, just look for UHD or 2160p copies.
- deleted 6y ago[deleted]
- im3w1l 6y ago50% is very impressive. It's not just a gold rush of low hanging fruit anymore, they did real work, created real benefits. I'm willing to pay a little tax on my devices or softwares for this.
- xiphias2 6y agoShouldn't deep learning based video codecs take over dedicated hardware video decoders as more tensor cores become available in all new hardware? NVIDIA's DLSS 2.0 supersampling is already moving into that direction.
- sp332 6y agoInstead of a video file or stream, that would be more like shipping a program that recreates the video. It might be cool, but it's not really feasible to play back that kind of thing on normal TV hardware.
- xiphias2 6y agoI'm not sure what you mean. There are already multiple research articles that show that deep neural network based video compression can be competitive, here's an example: https://papers.nips.cc/paper/9127-deep-generative-video-compression.pdf https://papers.nips.cc/paper/9127-deep-generative-video-comp...
- sitkack 6y ago> it's not really feasible How do you know this? TV hardware is on par with browsers. Anything is a program.
- somethingsome 6y agoIt's surprisingly difficult to guarantee a high quality on every kind of videos and formats using a neural network. Furthermore, the network should be able to handle all the corner cases (think about color profiles alone..)
- 0-_-0 6y agoQuestion is, how does it compare to AV1?
- miclill 6y agoI guess only time will tell. AV1 is supposed to be 30% better than HEVC and they claim H.266 is 50% better than HECV. This would mean that H.266 is roughly 30% better than AV1. By better I'm always referring to the bandwidth/space needed. But take this with more than a grain of salt since bandwidth/space are only one of many things that matter and also these comparisons are dependent on so many things like resolution, material (animatic/real), etc. etc.
- The_rationalist 6y agoAV1 is supposed to be 30% better than HEVC Source? If I recall correctly HEVC outperform AV1
- mda 6y agoI don't think your math adds up. Is 150 30% better than 130? It is only 14% better. Regardless, these early performance claims are most likely complete bullshit.
- occamrazor 6y agoIt’s about size: 50 is about 30% better than 70, which is 30% better than 100.
- TekMol 6y agoAt some point the compressed version of "Joker" will be 45 chars: "Sequel to Dark Night starring Joaquin Phoenix" Of course we will not have to film movies in the first place then. We will just put a description into a compressor start watching.
- prvc 6y ago8 MiB Shrek is kind of an AV1 meme at this point.
- mrfusion 6y agoSo how does it achieve this compression from a laypersons perspective?
- nickysielicki 6y agoI'm no expert when it comes to video codecs but I'm surprised that we're still able to see such strong claims of algorithmic improvements to h264, and now to h265. I'm also aware of how patent-encumbered this whole field is and I'm skeptical that this is just a money grab. This is really just a press release, what's actually new? Can it be implemented efficiently in hardware?
- bob1029 6y agoYour skepticism is very healthy, especially in this arena. With video codecs, information theory is ultimately the devil you must answer to at the end of the day. No amount of patents, specifications or algorithmic fantasy can get you away from fundamental constraints. It seems like the major trade-off being taken right now is along lines of using more memory to buffer additional frames. This can help you in certain scenarios, but in the general case, you cannot ever hope that a prior frame of video has any bearing on future frames of video. It is just exceedingly likely that most frames of video look much like prior frames. So, you can certainly play this game to a point, but you will quickly find yourself on the other end of the bell curve. You can also play games with ML, but I argue that you are going even further from the fundamental "truth" of your source data with this kind of technique, even if it appears to be a better aesthetic result in isolation of any other concern. There are also lots of one-off edge cases that have always been impossible to address with any interframe video compression scheme. Just look at the slowmo guys on youtube dump confetti on a 4K camera. No algorithm except for the dumbest intraframe techniques (i.e. JPEG) can faithfully reproduce scenes with information this dense, and usually at the expense of dramatic bandwidth increases. Bandwidth is cheap and ubiquitous. I say we just use the algorithms that are the fastest and most efficient for our devices. We aren't in 2010 sucking 3G or edge through a straw anymore. Most people can get 20+mbps in their smartphones in decently-populated areas.
- freeone3000 6y agoThe advantage of the H-series of codecs is strong support of hardware implementation. This has been a selling point since H.262. You can get a H.265 IP core from Xilinx, Intel, and other major vendors -- so the actual runtime performance for H.266 (once a core is available) will be very low and constant (and comparable to current codecs). Bandwidth and storage space are real costs, despite the handwaving around it, and reducing these requirements while not reducing visual quality is an important step. As for "information-dense scenes": Pathologic cases such as the HBO intro screen are encoded into modern codecs as noise, and regenerated client-side, because there's no actual information there. These scenes are either engineered or pure noise.
- shmerl 6y agoAnother patent encumbered monstrosity? No, thanks. Enough of this junk. Some just never learn.
- baybal2 6y agoH.265 is still not mainstream, and not used to full extend of its performance I'm not sure if 265 is worth spending efforts on now when 266 is about to crash the party, and will be equally adopted at least "equally poorly"
- bufferoverflow 6y agoH.266 will take many years to become a usable standard. H.265 is quite mainstream - even some cheap smartphones shoot it. Many modern DLSRs / mirrorless cameras shoot it.
- crazygringo 6y agoH.265 seems pretty mainstream by now. Older devices obviously don't support it in hardware, but pretty much all newer ones seem to, no? It's just a slow percolation throughout the ecosystem as people buy new hardware and video servers selectively send the next-generation streams to those users. The effort on h.265 has already been spent. Now it looks like h.266 is the next generation. It's going to be years before chips for it will be in devices. That's just how each new generation works.
- ConsiderCrying 6y agoH.265 seems to be gaining traction slower because many of the older devices, including laptops and some smart TVs don't support it. H.264 became ubiquitous for piracy since it offered tiny size and worked on older devices, making it the perfect choice for those in poorer countries where tech isn't the first priority in a household. I wonder if H.266 will run into the same problems as H.266.
- anordal 6y agoNobody mentioning EVC? Worth a read for anyone concerned about patent licensing: https://en.wikipedia.org/wiki/Essential_Video_Coding https://en.wikipedia.org/wiki/Essential_Video_Coding There are 3 video coding formats expected out of (former) MPEG this year: https://www.streamingmedia.com/Articles/Editorial/Featured-Articles/Inside-MPEGs-Ambitious-Plan-to-Launch-3-Video-Codecs-in-2020-134694.aspx?utm_source=related_articles&utm_medium=gutenberg&utm_campaign=editors_selection https://www.streamingmedia.com/Articles/Editorial/Featured-A... So this isn't necessarily the successor to HEVC (except that it is, in terms of development and licensing methods).
- robomartin 6y agoAmong other things, I have worked with and developed technology in the uncompressed professional imaging domain for decades. One of the things I always watch out for is precisely the terminology and language used in this release: "for equal perceptual quality" Put a different way: We can fool your eyes/brain into thinking you are looking at the same images. For most consumer use cases where the objective is to view images --rather than process them-- this is fine. The human vision system (HVS, eyes + brain processing) is tolerant of and can handle lots of missing or distorted data. However, the minute you get into having to process the images in hardware or software things can change radically. Take, as an example, color sub-sampling. You start with a camera with three distinct sensors. Each sensor has a full frame color filter. They are optically coupled to see the same image through a prism. This means you sample the red, green and blue portions of the visible spectrum at full spatial resolution. If we are talking about a 1K x 1K image, you are capturing one million pixels of each, red, green and blue. BTW, I am using "1K" to mean one thousand, not 1024. Such a camera is very expensive and impractical for consumer applications. Enter the Bayer filter [0]. You can now use a single sensor to capture all three color components. However, instead of having one million samples for each components you have 250K red, 500K green and 250K blue. Still a million samples total (that's the resolution of the sensor) yet you've sliced it up into three components. This can be reconstructed into full one million samples per color components through various techniques, one of them being the use of polyphase FIR (Finite Impulse Response) filters looking across a range of samples. Generally speaking, the wider the filter the better the results, however, you'll always have issues around the edges of the image. There are also more sophisticated solutions that apply FIR filters diagonally as well as temporally (use multiple frames). You are essentially trying to reconstruct the original image by guessing or calculating the missing samples. By doing so you introduce spatial (and even temporal) frequency domain issues that would not have been present in the case of a fully sampled (3 sensor) capture system. In a typical transmission chain the reconstructed RGB data is eventually encoded into the YCbCr color space [1]. I think of this as the first step in the perceptual "let's see what we can get away with" encoding process. YCbCr is about what the HVS sees. "Y" is the "luma", or intensity component. "Cb" and "Cr" are color difference samples for blue and red. However, it doesn't stop there. The next step is to, again, subsample some of it in order to reduce data for encoding, compression, storage and transmission. This is where you get into the concept of chroma subsampling [2] and terminology such as 4:4:4, 4:2:2, etc. Here, again, we reduce data by throwing away (not quite) color information. It turns out your brain can deal with irregularities in color far more so than in the luma, or intensity, portion of an image. And so, "4:4:4" means we take every sample of the YCbCr encoded image, while "4:2:2" means we cut down Cb and Cr in half. There's an additional step which encodes the image in a nonlinear fashion, which, again, is a perceptual trick. This introduces Y' (Y prime) as "luminance" rather than "luma". It turns out that your HVS is far more sensitive to minute detail in the low-lights (the darker portions of the image, say, from 50% down to black) than in the highlights. You can have massive errors in the highlights and your HVS just won't see them, particularly if things are blended through wide FIR filters during display. [3] Throughout this chain of optical and mathematical wrangling you are highly dependent on the accuracy of each step in the process. How much distortion is introduced depends on a range of factors, not the least of which is the way math is done in software or chips that touch every single sample's data. With so much math in the processing chain you have to be extremely careful about not introducing errors by truncation or rounding. We then introduce compression algorithms. In the case of motion video they will typically compress a reference frame as a still and then encode the difference with respect to that frame for subsequent frames. They divide an image into blocks of pixels and then spatially process these blocks to develop a dictionary of blocks to store, transmit, etc. The key technology in compression is the Discrete Cosine Transform (DCT) [4]. This bit of math transforms the image from the spatial domain to the frequency domain. Once again, we are trying to trick the eye. Reduce information the HVS might not perceive. We are not as sensitive to detail, which means it's safe to remove some detail. That's what DCT is about. So, we started with a 3 sensor full-sampling camera, reduced it to a single sensor and three away 75% of red samples, 50% of green samples and 75% of blue samples. We then reconstruct the full RGB data mathematically, perceptually encode it to YCbCr, apply gamma encoding if necessary, apply DCT to reduce high frequency information based on agreed-upon perceptual thresholds and then store and transmit the final result. For display on an RGB display we reverse the process. Errors are introduced every step of the way, the hope and objective being to trick the HVS into seeing an acceptable image. All of this is great for watching a movie or a TikTok video. However, when you work in machine vision or any domain that requires high quality image data, the issues with the processing chain presented above can introduce problems with consequences ranging from the introduction of errors (Was that a truck in front of our self driving car or something else?) to making it impossible to make valid use of the images (Is that a tumor or healthy tissue?). While H.266 sounds fantastic for TikTok or Netflix, I fear that the constant effort to find creative ways to trick the HVS might introduce issues in machine vision, machine learning and AI that most in the field will not realize. Unless someone has a reasonable depth of expertise in imaging they might very well assume the technology they are using is perfectly adequate for the task. Imagine developing a training data set consisting of millions of images without understanding the images have "processing damage" because of the way they were acquired and processed before they even saw their first learning algorithm. Having worked in this field for quite some time --not many people take a 20x magnifying lens to pixels on a display to see what the processing is doing to the image-- I am concerned about the divergence between HVS trickery, which, again, is fine for TikTok and Netflix and MV/ML/AI. A while ago there was a discussion on HN about ML misclassification of people of color. While I haven't looked into this in detail, I am convinced, based on experience, that the numerical HVS trickery I describe above has something to do with this problem. If you train models with distorted data you have to expect errors in classification. As they say, garbage-in, garbage-out. Nothing wrong with H.266, it sounds fantastic. However, I think MV/ML/AI practitioners need to be deeply aware of what data they are working with and how it got to their neural network. It is for this reason that we've avoided using off-the-shelf image processing chips to the extent possible. When you use an FPGA to process images with your own processing chain you are in control of what happens to every single pixel's data and, more importantly, you can qualify and quantify any errors that might be introduced in the chain. [0] https://en.wikipedia.org/wiki/Bayer_filter https://en.wikipedia.org/wiki/Bayer_filter [1] https://en.wikipedia.org/wiki/YCbCr https://en.wikipedia.org/wiki/YCbCr [2] https://en.wikipedia.org/wiki/Chroma_subsampling https://en.wikipedia.org/wiki/Chroma_subsampling [3] https://en.wikipedia.org/wiki/Gamma_correction https://en.wikipedia.org/wiki/Gamma_correction [4] https://www.youtube.com/watch?v=P7abyWT4dss https://www.youtube.com/watch?v=P7abyWT4dss
- xvilka 6y agoDue to the patenting Fraunhofer probably did more harm to humanity than something good. At least its software division.
- pxf 6y agoWill it be used? Probably the last one that does not use some sort of AI Compression.See this for image compression https://hific.github.io/ https://hific.github.io/ In the next 10 years AI Compression will be everywhere. The problem will be standartisation. Classic compression algoritms can't beat AI ones.
- crazygringo 6y agoAI compression is super, super cool... but while standardization is certainly a major issue, isn't the model size a much larger one? Given that model sizes for decoding seem like they'll be on the order of many gigabytes, it will be impossible to run AI decompression in software, but will need chips, and chips that are a lot more complex (expensive?) than today's. I think AI compression has a good chance of coming eventually, but in 10 years it will still be in research labs. There is absolutely no way it will have made it into consumer chips by then.
- pxf 6y ago"Isn't the model size a much larger one?" yap It will probably be different, and systems will have to download the weights and network model, as new models come in, I don't think that we will have a fixed model with fixed weights, the evolution is too fast. Decoding will take place using the AI chip on the device aka "AI accelerator"
- jonpurdy 6y agoMy 2012 Mac Mini has quickly become much less useful since YouTube switched from H.264 (AVC) to VP9 for videos larger than 1080p a couple of years ago (Apple devices have hardware decoders). I've tested 4K h.264 videos and they play wonderfully thanks to the hardware. My internet connection speeds and hard drive space have increased much faster than my CPU speeds (internet being basically a free upgrade). So I don't appreciate new codecs coming out and obsoleting my hardware to save companies a few cents on bandwidth. H.264 got a good run in, but there isn't a "universal" replacement for it where I can buy hardware with decoding support that will work for at least 5-10 years.
- crazygringo 6y agoHonestly, the expectation that a 2012 computer will play 4K video seems a little unreasonable, no? 4K video virtually didn't even exist back then. I'm actually amazed it even handles it in h.264. This isn't about saving companies a few cents on bandwidth. It's about halving internet traffic, about doubling the number of videos you can store on your phone. That's pretty huge. You can still get h.264 video in 1080p on YouTube so your computer is still meeting the expectations it was manufactured for.
- jonpurdy 6y agoIt's not so much about it being able to handle 4K, but Youtube already has too low of a bitrate for 1080p (resulting in MPEG artifacts, color banding, etc). So I like to watch in 2.7K or 4K downsampled, since at least I get a higher bitrate. The bigger problem that I didn't mention is with videoconferencing: FaceTime is hardware accelerated and has no issues with 720p, but anything WebRTC seems to prefer VP8 or VP9 codecs, which fails on my Mini and strains my 2015 MBP. Feels like a waste of perfectly good hardware to me.
- rahimnathwani 6y agoMaybe you can force the YouTube web app to use the H.264 version? https://github.com/erkserkserks/h264ify https://github.com/erkserkserks/h264ify Or is there no longer an H.264 4k version?
- crazygringo 6y agoThey talk about saving 50% of bits over h.265, but also talk about it being designed especially for 4K/8K video. Are normal 1080p videos going to see this fabled 50% savings over h.265? Or is the 50% only for 4K/8K, while 1080p gets maybe only 10-20% savings? The press release unfortunately seems rather ambiguous about this.
- znpy 6y agoI wonder how small would one of those 700mb divx/xvid movies would be if compressed with this new encoding method.
- prvc 6y agoIt is mildly amusing that the very simple vector art "VVC" logo on their webpage is displayed by sending the viewer a 711 KB .jpg file.
- superkuh 6y agoI'd rather have slightly larger files that don't take hardware acceleration only available on modern CPU to decode without dying (ie, h264). Streaming is creating incentives for bad video codecs that only do one thing well: stream. Other aspects are neglected. And it's not like any actual 4K content (besides porn, real, nature, or otherwise) actually exists. Broadcast and movie media is done in 2K then extrapolated and scaled to "4K" for streaming services.
- crazygringo 6y agoHuh? TV and movies are widely shot with 4K cameras these days. What is 2K? I've never even heard of a "2K" camera. Where did you get the idea things are being filmed in "2K" and being scaled to 4K? Genuinely curious where you're getting this information from. Or are you confused because 1080p refers to the vertical resolution while 4K refers to the horizontal resolution?
- superkuh 6y agohttps://www.engadget.com/2019-06-19-upscaled-uhd-4k-digital-intermediate-explainer.html https://www.engadget.com/2019-06-19-upscaled-uhd-4k-digital-... is one easily found example but it wasn't where I had read it. I'm pretty sure I've seen it on HN itself. edit: here's another https://old.reddit.com/r/cordcutters/comments/9x3v4e/just_learned_most_4k_content_is_actually_2k/ https://old.reddit.com/r/cordcutters/comments/9x3v4e/just_le...
- crazygringo 6y agoOK, so by 2K you mean 1080p. That's a very unusual nomenclature but I see what you mean, thanks. The top link in the reddit thread disproves what you're saying though: https://4kmedia.org/real-or-fake-4k/ https://4kmedia.org/real-or-fake-4k/ Somewhere between a third and a half of films are listed as "real 4K". So there is actually tons of real 4K content. (And the list is just films -- there are plenty of streaming TV shows in real 4K too, like Mrs Maisel.) There might be another reason for the misperception -- it's true that film editing is generally done in something lower-quality like compressed 1080p, but that's just for speed/space while you work. All the clips "point" to the 4K originals, so when the final master is produced, it's still produced out of that "real" 4K.
- irrational 6y agoAnyone know how this compares to AV1?
- qwerty456127 6y agoHow many weeks does it take to encode a 1-minute video on an average (non-gaming, I mean without a huge fancy GPU card or an i9/Threadripper CPU) PC?
- blacklion 6y agoWhat I don't understand, why do internationa; standardization organizations allows patent-encumbered technologies to become de-jure standards. MPEG, WiFi, GSM… IMHO, intentional standards must be implementable without any patent fees, or they are very bad standards.
- freeone3000 6y agoThere's no law requiring wifi - "de facto". And they're standards because they're quite good! They have hardware support and parallelization and account for all use cases, even the marginal ones, and have reference implementations and support. Standards orgs don't care about patents because they're not relevant. This isn't a case of trolling - this is literally a software patent being used for its intended purpose by its developer, to extract profit by coming up with a new idea, and letting others use it.
- ajnin 6y agoInternational standardization organizations (like ISO) are not governmental organizations. They are private entities, which sometimes become too involved in official standards. But they are controlled by whoever funds them.
- holloway 6y agoA quote from an Ecma presentation "ECMA for instance has made all the standards for DVD and optical disks. There were 5 recording formats. So there you are a little bit uneasy, of course. And again after a few beers I can ask the people in the room. Why do you want to have 5 formats? Do you still call that standardization? The answer is always the same: You are well paid. Shut up" https://youtu.be/wITyO71Et6g?t=226 https://youtu.be/wITyO71Et6g?t=226
- blibble 6y agoH.265 went absolutely nowhere
- lmm 6y agoStandards organisations are older than publicly-available software. The concept of "reasonable and non-discriminatory" patent licensing was what they went with, and it seemed sensible at a time when goods were physical and the idea of giving away a product for which standards would be relevant would be ludicrous.
- Havoc 6y agoThe fact that the underground scene is still pumping 264 instead of 265 (I'd estimate 90/10 split optimistically) tells me the real world is not quite ready for 266. So I guess it comes down to 266 hw support. Or powerful CPUs that can push sw decoding?
- liquid153 6y agoWill devices need new hardware. Also I thought companies were all on board with royalty free VP9
- charliebrownau 6y agoWhat happened to AV1 ?