3 ms·
As others have replied, this is reasonable general feedback, but in this specific case the work was done carefully. Table 1 from the linked paper (https://arxi
by boulos 2y ago
As others have replied, this is reasonable general feedback, but in this specific case the work was done carefully. Table 1 from the linked paper (https://arxiv.org/pdf/2411.05007 https://arxiv.org/pdf/2411.05007) includes a variety of metrics, while an entire appendix is dedicated to quality comparisons.
By showing their work side-by-side with other quantization schemes, you can also see a great example of the flavor of different results you can get with these slight tweaks (e.g., ViDiT INT8) and that their quantization does a much better job in reproducing the "original" (Figure 15).
In this application, it's not strictly true that you care to have the same results, but this work does a pretty good job of it.
- djoldman 2y agoAgreed. Once a model has been trained, I believe the main metrics people care about are 1. inference speed 2. memory requirements 3. quality of output. There are usually tradeoffs here. Generally you get a lower memory requirement (a good thing), sometimes faster inference (a good thing), but usually a lower quality of output. I don't think reproduction of original output is the typical goal.