4 ms·
Worth mentioning: their demo only uses two passes for progressive AVIF -- it's just their choice. You can have up to two more intermediate passes, so at 30% you
by juliobbv 13d ago
Worth mentioning: their demo only uses two passes for progressive AVIF -- it's just their choice. You can have up to two more intermediate passes, so at 30% you can have something much closer to the original image.
Even still, the demo proves that AVIF can deliver a usable image with up to 3x as fewer bytes as JXL!
- bilkow 13d ago> their demo only uses two passes for progressive AVIF -- it's just their choice Also worth mentioning: AFAIK "their choice" here is just the default, i.e. what `avifenc --progressive` outputs, so maybe if the default is suboptimal, it could be improved? > Even still, the demo proves that AVIF can deliver a usable image with up to 3x as fewer bytes as JXL! Usable for what? As a clearly loading image / placeholder, I personally like the JXL version better (as already indicated in my previous comment). As a final image I think both of them are unusable, and if that's the intention I think it would be better to encode both aiming for very low quality, maybe also reducing the resolution, without progressive loading, then compare. My understanding is that AVIF is usually better at very low bitrates, so it would probably be better, but I don't think this example "proves" that.
- juliobbv 13d ago> AFAIK "their choice" here is just the default Well, the default in avifenc can always be changed. Do keep in mind there's no "one size fits all" implementation, as customers desire different loading tradeoffs. You might be surprised, but during testing (outside HN), we've seen people actually prefer "2 layer" loading. I'm surprised HN likes progressive loading to be more granular, and use that to push back. I'm wondering if there are generational differences at play? Maybe it's the difference of being used to the blurhash vs. old-school JPEG loading experience. Anyway, the more expressive mode in avifenc is `--layered` (yes, I know the name is weird). > Usable for what? Well, focusing on the "Poke bowl" example: with AVIF you can clearly tell apart each ingredient at 8kB. You can derive enough semantic understanding just from this base layer. Having decent edge preservation helps significantly here. On the other hand, JXL is just too blurry at 8kB to make sense of the image -- only the egg and carrot are recognizable, maaybe the cucumber? IMO JXL needs the pass at ~28kB to make everything salient, including sprouts and beet. Yes, I know there are subjective effects at play and I'm sure we'll disagree on exact image thresholds, but recognizing objects within an image is so important in real-life use cases.
- bilkow 13d ago> Well, the default in avifenc can always be changed. Sure, that was kind of my point, but I think it's important to point out it's the default. I had read somewhere else in the comments that they just "happened" to use 2 passes (that I now realized was also written by you), that I interpreted as them maybe making a bad decision in the comparison page, but I think it's very reasonable to use the default. I also usually prefer to use the default unless I have a good reason not to do so. And I think that `--layered` is fine as a name, and I had realized that as the docs were right below the docs for `--progressive`, which do mention it encodes a "layered image", although it could use an example on how to use it effectively. > I'm surprised HN likes progressive loading to be more granular, and use that to push back. Not necessarily? Using more passes would probably help in specific the metric in the comment I was responding to, by making AVIF possibly look better on a larger fraction of the decoding time. I mostly interpreted your comment as saying that it would have been better if they used more passes, and that's on me. The reason I was comparing the breakpoints in the decode time was because I was responding to an allegation that "even there, AVIF looks better for 90% of the decode time", and I decided to point out that's not not the case and chose some pretty clear breakpoints that are very hard to argue about. Very subjectively the actual breakpoints could be a lot earlier, for the Poke the AVIF only looks better (in the sense that you can figure out what's there better) in a 6% range (from 2 to 8%) or maybe 9% (from 2% to 11%) of the decode time. There are many (relatively big) seeds that are just not there on the AVIF side, but you can distinguish on the JXL side, even if they're blurry. Percentages used in reference to the JXL side. > I'm wondering if there are generational differences at play? Maybe it's the difference of being used to the blurhash vs. old-school JPEG loading experience. Hmmm I don't know, maybe. TBF with today's usual internet speeds progressive loading isn't as important as it was as more steps make less sense if you have a good connection. > Well, focusing on the "Poke bowl" example: with AVIF you can clearly tell apart each ingredient at 8kB. You can derive enough semantic understanding just from this base layer. Having decent edge preservation helps significantly here. Sure, I see what you mean, you can distinguish more image element in the AVIF example, that's true. I don't like the lack of consistency and some of the artifacts to the point where I'd prefer not using progressive encoding, so that's not what I'd call "usable", but that seems like a matter of taste. It surely does look closer to a "final image" than the JXL at that point. As a note, I think what I'd prefer would be 2 passes, with the 1st pass being around how JXL looks at around 11% / 39 kB. You can get a rough idea of what's the image about and the elements on it, with roughly the right colors, but it's still clearly loading, so there's no confusion, and without many "artifacts" like the progressive AVIF in the example (no idea if AVIF is able to make progressive encoding in a way that's more "blurry"). The caveat is that more steps are better when it takes too long (more than a few seconds) so the user knows it's not stuck.
- innocent_name 13d ago>Even still, the demo proves that AVIF can deliver a usable image with up to 3x as fewer bytes as JXL! I was taught to be skeptical of objective benchmarks; Why should we trust them when people are benchmaxxing? https://habr.com/en/articles/700726/ https://habr.com/en/articles/700726/
- juliobbv 13d agoI'm not following.... the demo involves you look at images as they get decoded. Objective benchmarks are beside the point here.
- innocent_name 13d agoOh, i left the wrong quotation whereas i intended to reply to another message. Throughout the blogpost you're quality matching and comparing encoders based on objective metrics whereas it'd be more telling to get the crowd subjective comparisons. I think it's pretty evident that most codecs benchmaxx to the point of objective metrics being useless.
- juliobbv 13d agoAh, no worries! I can't speak for Iris and Aperture, but both "tune IQ" modes in SVT-AV1 and libaom had extensive human evaluations to make sure they weren't accidentally being benchmaxxed at the expense of subjective quality.