5 ms·
>GPT-3.5 has 175B parameters versus 70B parameters in Llama 2 We know that for the original version of GPT-3.5, but my assumption was that Turbo was a distille
by ImprobableTruth 3y ago
>GPT-3.5 has 175B parameters versus 70B parameters in Llama 2
We know that for the original version of GPT-3.5, but my assumption was that Turbo was a distilled smaller model (which is why it uses OAI's new vocab & is so much faster).
If that's not the case, what could be the explanation for it being faster?
- rasbt 3y agoI think so too. But in general, it could also be due to other reasons: faster hardware, lower timeout for batched inference, optimizations like flash attention and flash attention 2, quantization, ... I'd say that it's probably a mix of all of the above (incl some distillation).
- sebzim4500 3y agoIt is widely believed that GPT-3.5 is a MoE model, which means it could have 175B parameters but still be much lower latency than GPT-3
- rasbt 3y agoInteresting, I thought GPT-3.5 was considered GPT-3 + InstructGPT-style RLHF on a large scale, whereas GPT-4 is considered to be an MoE model.
- rgbrgb 3y agoWhy would MoE make it lower latency?
- sebzim4500 3y agoIt's easier to parallelize so you can throw more GPUs at a single request (or really, batch of requests)
- rgbrgb 3y agoInteresting, yeah I buy that, thanks. Building my intuition with this stuff. Anyone seen a good open-source implementation of MoE with Llama yet?
- sebzim4500 3y agoYou can't just turn an existing model into MoE, they need to be trained from scratch unfortunately. I'm not aware of any open source MoE models, they are complicated and probably not that useful if you want to run them on your own hardware.
- Me1000 3y agoWould you mind correcting my misunderstanding here? Code Llama is a fine tuned version of Llama2 (i.e. not trained from scratch). If I fine tuned Llama2 with a bunch of law text and had Law Llama, and fined tuned a couple more with some history text and science text. Why wouldn't Code Llama, Law Llama, History Llama, and Science Llama not be the experts in my MoE setup? Seems like I just need a simple router in front of those models to direct the prompt to the right expert.
- sebzim4500 3y agoThat could work, but I'd expect the following issues: * For a lot of prompts every fine tuned model will make the same mistakes (they mostly share the same weights after all) and so you aren't getting nearly as much benefit as e.g. GPT-4 gets. * It's going to be really expensive at inference time, since you have to run multiple models even though in most cases they won't help much * Normally when people talk about hobbyists doing finetuning they mean <1M tokens, whereas Code Llama Python was finetuned on 100B tokens, way outside most people's price range. For the finetuning that you can afford, you can't teach it new knowledge, just show it how to apply the knowledge it already has.
- phillipcarter 3y agoI think that unless (until?) OpenAI releases information about the model itself and the inference engine it runs on, everything is just speculation. Clearly, there's impressive ML and systems engineering at play with GPT-3.5-turbo given how capable, fast, and scalable to their customer base it is.
- visarga 3y agoThere is also speculative sampling - you decode a few tokens with a smaller model, then use the big model to validate them simultaneously. The big model might trim the prediction up to a point and add an extra token. Then cycle again with the small model -> 2-2.5x speedup