6 ms·
> "We then train several models from 400M to 15B on the same pre-training mixture for up to 1 × 1022 FLOPs." Seems that for the last year or so these models ar
by techbruv 3y ago
> "We then train several models from 400M to 15B on the same pre-training mixture for up to 1 × 1022 FLOPs."
Seems that for the last year or so these models are getting smaller. I would be surprised if GPT-4 had > the number of parameters as GPT-3 (i.e. 175B).
Edit: Seems those numbers are just for their scaling laws study. They don't explicitly say the size of PaLM 2-L, but they do say "The largest model in the PaLM 2 family, PaLM 2-L, is significantly smaller than the largest PaLM model but uses more training compute.". So likely on the range of 10B - 100B.
- thewataccount 3y agoI've heard Bard was previously 3B parameters but I could never find a good source for it. I honestly think the end game here is running on consumer devices, 7B and under need ~4GB of ram to actually run which is likely the max reasonable requirement for consumer devices. That said medium end hardware can do 15B, anything larger then this is currently something only "enthusiasts" can run. If it is small enough to run on consumer devices then they don't have to pay for the inference compute at that point, and presumably the latency will be improved for consumers.
- int_19h 3y agoThe current state of consumer devices isn't static, either, and existing hardware (even GPU) is suboptimal for the current crop of LLMs - it does way more than it actually needs to do.
- famouswaffles 3y agoThose are the numbers for the scaling law tests they did. Not necessarily Palm 2 range.
- tempusalaria 3y agoGPT-4 is way slower than GPT-3. Unless they are artificially spiking the latency to hide parameter count, it’s likely around 1trn params
- thewataccount 3y agoYeah 1 to 2 trillion is the estimates I've heard. Given the 25 messages / 3 hour limit in chatGPT, I don't think they've found a way to make it cheap to run.
- tempusalaria 3y agoYep. I’m guessing PaLM 2 is about 200bln params as it seems clearly stronger than chinchilla
- dontupvoteme 3y ago1. there's no reason to think OpenAI wouldn't also be going the artificial scarcity route as have so many other companies in the past 2. Microsoft may not like them using too much azure compute and tell them to step off. Rumor has it they're trying to migrate github to it and it's seemingly not going ideal. And they're certainly nothing more than another microsoft purchase at this point.
- akiselev 3y agoOpenAI has a 40k token per minute rate limit on their GPT4 API too so I doubt it's artificial scarcity.
- dontupvoteme 3y agoPerhaps. I found it was far too easy to hit the API limit with their old codex models, though that may have been limited to a small GPU cluster given it was pretty obscure compared to chatgpt and even davinci.
- thewataccount 3y agoBased on GPT3.5 supposedly using 8x A100's per query and the suspected magnitude size difference with GTP4 I really think they're struggling to run it. At this stage I think they'd have more to benefit by making it more accessible, there's several use cases I have (or where I work) that only really make sense with GPT4, and it's way too expensive to even consider. Also AFAIK Github Copilot is still not using GPT4 or even a bigger CODEX, and GPT4 still outperforms it especially in consistency (I'm in their copilot chat beta).
- gwern 3y agoFor 'Palm-2', read, 'T5-2'.