100 ms·
Isn't it the case that we literally have no clue how GPT4 and GPT3.5 are different in terms of training, given OpenAI doesn't want to disclose anything at all?
by feanaro 4y ago
Isn't it the case that we literally have no clue how GPT4 and GPT3.5 are different in terms of training, given OpenAI doesn't want to disclose anything at all?
- computerex 4y agoWe don' have the details, it is true. But empirically and based on their report gpt-4 is notably better than chatgpt.
- feanaro 4y agoBetter, yes, and for that we have evidence. But is the improvement stemming simply from even more data? That's what I'm questioning.
- stevenhuang 4y agoIt's speculated it has same number of parameters, but more compute and is multi modal.
- computerex 4y agoThis paper is pretty approachable and goes over the "scaling laws" in detail: https://arxiv.org/abs/2206.07682 https://arxiv.org/abs/2206.07682 In short, yes. More data, higher quality data, more epochs on the data. That is the name of the game.
- feanaro 4y agoThat paper doesn't discuss GPT-4 at all. It does however contain this interesting excerpt (emphasis mine): > Although we may observe an emergent ability to occur at a certain scale, it is possible that the ability could be later achieved at a smaller scale—in other words, model scale is not the singular factor for unlocking an emergent ability. As the science of training large language models progresses, certain abilities may be unlocked for smaller models with new architectures, higher-quality data, or improved training procedures. For example, there are 14 BIG-Bench tasks5 for which LaMDA 137B and GPT-3 175B models perform at near-random, but PaLM 62B in fact achieves above-random performance, despite having fewer model parameters and training FLOPs. So it's not obvious that it should be so straightforward.
- typon 4y agoIt's not true we know nothing. We know a little bit by using the two models from their API. Given the time per inference and the limit on messages per day for GPT4, I'm willing to bet it's doing around 10x more compute than GPT3.5. If that's because it has 10x more weights, I don't know. But it wouldn't be a terrible guess.
- feanaro 4y agoSo your estimate is that GPT4 has 1.75 trillion weights?
- dwaltrip 4y agoIs there anything that affects inference compute time besides the number of parameters? Assuming same hardware, etc.
- typon 4y agoYes - for example adding memory to the attention mechanism (similar to RETRO or Memorizing Transformers paper)