6 ms·
early int4 experiments seem to indicate it's possible but you do lose performance, see this thread https://www.reddit.com/r/MachineLearning/comments/11i4olx/d_i
by v64 4y ago
early int4 experiments seem to indicate it's possible but you do lose performance, see this thread https://www.reddit.com/r/MachineLearning/comments/11i4olx/d_is_it_possible_to_run_metas_llama_65b_model_on/ https://www.reddit.com/r/MachineLearning/comments/11i4olx/d_...
edit: to clarify, it may be possible to get this loss back and there is reason to be optimistic
- CuriouslyC 4y agoProbably the best method is to just train it on int4 in the first place. Fine tuning after quantization would definitely help though.
- sp332 4y agoIsn't that backwards? You need fairly good resolution during training or your gradients will be pointing all over the place. Once you've found a good minimum point, moving a little away from it with reduced precision is probably OK.
- brookst 4y agoI have no idea what the right answer is, but I think the argument for int4 training is that the loss measurements would take the lower resolution of the model as a whole into account. Is it better to have billions of high resolution parameters and quantize them at the end, or to train low resolution parameters where the training algorithms see the lower resolution? It’s beyond me, but I’d love to know.
- Scene_Cast2 4y agoBut by default, training algos don't see the lower resolution, your gradient just doesn't work as well. There is a body of research on how to make training aware of / adapt to the lower precision.
- bick_nyers 4y agoI think the answer is it depends, and further, a dynamic approach may be best. Imagine you are going on a hike, and you have different maps at various resolutions (levels of detail). When planning the hike, you will want to see the zoomed out picture to get general directions, elevations and landmarks identified. Then you can zoom in to see the actual trails themselves, to identify your route, and then you zoom in even further when you are on the ground actually walking, avoiding obstacles along the way. Different resolutions draw your attention to different types of features.
- rfoo 4y agoGP could be mentioning quantization aware training, during which the weight and gradient are still computed in fp16/fp32.
- CuriouslyC 4y agoIt can go farther than that, it seems like the weight gradients are the main thing where the precision is a bottleneck (see https://arxiv.org/abs/1805.11046 https://arxiv.org/abs/1805.11046).
- nl 4y ago> Probably the best method is to just train it on int4 in the first place Unclear why you think that since experiments show the opposite. In general the gradient seems to get too "bumpy" to do good gradient decent at lower levels of precision. There are some papers showing that making the training loop aware of quantitization can help ultimate quantizied performance but I'm not aware of this being implemented at large scale.
- bick_nyers 4y agoWhat if you smooth the gradient, either by interpolating/removing data points that make the surface "jagged", or maybe change the "window" of gradient descent, meaning instead of using a tangent (derivative, infinitesimally small window) you use a secant (???, window of specified length, likely calculated from the data space). Forgive my lack of proper terminology here.
- nl 4y agoSure, there are multiple ways to reduce the complexity of your loss-space, but the issue is that you usually want these small gradient values because they are important. Roughly if you "smooth over what appears to be a small hole" often you'll miss a large space that needs to be explored (obviously this is multi-dimensional but you get the idea). However you can reduce memory by doing mixed-precision training if you are careful. See section "2.3.1. Loss Scaling To Preserve Small Gradient Magnitudes" in https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html https://docs.nvidia.com/deeplearning/performance/mixed-preci...
- bick_nyers 4y agoSo then you would need to do some kind of mesh simplification that also preserves the topology, that makes sense. I'm not quite sure I understand what they are describing in 2.3.1, are they scaling those small gradient magnitudes larger to try to "pull" you into those holes faster? I was thinking the a way to go about it would be to just increase the "mesh resolution" near the small hole, which in this case would be use a larger precision in the area local to the hole.