4 ms·
You're definitely right, the quantization function and its values definitely have an impact on performance. For 1 bit I think I tried something like -1/+1, -.5
by maxlam 9y ago
You're definitely right, the quantization function and its values definitely have an impact on performance.
For 1 bit I think I tried something like -1/+1, -.5/+.5, -.25/+.25, -.333/+.333. and something like -10/+10 -- (and I think a few more). It seemed -.333/+.333 worked the best while +10/-10 did the worst on the google analogy task (getting like 0% right). All this was tuned on 100MB of Wikipedia data.
- yorwba 9y agoHave you considered doing gradient descent on the quantization steps? It looks to me like the model should be differentiable with respect to those values, so I'm not sure why you'd have to fix them to a constant.
- maxlam 9y agoHm what do you mean? I'm not quite seeing how to differentiate with respect to the quantization steps.
- yorwba 9y agoSay you have a function f(q(x)) where q quantizes x into one of s_1, ..., s_n. Then if q(x) = s_i for a certain x, df/ds_i = df/dq and df/ds_j = 0 for all j != i. That breaks down for values of x precisely at the boundary between steps, so I should have qualified "differentiable" with "almost everywhere". It also occurs to me that this might interact strangely with the approximation dq/dx = 1, but since the quantization steps are globally shared, I think it should be stable anyway. If the evaluation suite for your code doesn't require too much manual interaction, I might try and see for myself.
- maxlam 9y agoThat's definitely an interesting idea -- it seems this would allow for boundaries that "change" along with the data (instead of having static boundaries as it is). Would be interested to know how that turns out!