3 ms·
Is Best-of-N Sampling standard practice these days in Inference? Sounds expensive on the face of it. I am surprised because I thought the trend was towards chea
by justanotheratom 1y ago
Is Best-of-N Sampling standard practice these days in Inference? Sounds expensive on the face of it. I am surprised because I thought the trend was towards cheaper inference.
- diwank 1y agoFor reasoning models, this would actually improve exploration efficiency and hence possibly allow higher performance for the same compute budget. As in, if you want to sample from multiple rollouts for the same prompt, it's more efficient if the model is able to produce diverse thought directions and consider them to find the best response as opposed to going down similar trajectories and waste compute.
- codelion 1y agoNot standard but one of several techniques, you can see them in our open source inference proxy - https://github.com/codelion/optillm https://github.com/codelion/optillm Cerebras has used optillm for optimising inference with techniques like CePO and LongCePO.
- peepeepoopoo114 1y agoAlmost all of the efficiency gains have come from shedding bit precision, but the problem is that AI labs are now running out of bits to shed. The move to reduced precision inference has been masking the insane unsustainability of compute scaling as a model improvement paradigm.
- nullc 1y agoIs there really a limit on bits to shed? I suspect not. Take N gates, normalize them, represent them as points on the surface of a hypersphere. Quantize the hypersphere as coarsely as you need to get the precision you want. Want less precision but your quantization is getting too coarse? Increase N. Fast algebraic codes exist to convert positions on a hyperspheric-ish surfaces to indexes and vice versa. Perhaps spherical VQ isn't ideal-- though I suspect it is, since groups of weights often act as rotations naturally-- but some other geometry should be good if not.