3 ms·
Original PaLM was 540B so significantly smaller could mean anything from 350B down really
by tempusalaria 3y ago
Original PaLM was 540B so significantly smaller could mean anything from 350B down really
- espadrine 3y agoI tried my hand at estimating their parameter count from extrapolating their LAMBADA figures, assuming they all trained on Chinchilla law: https://pbs.twimg.com/media/Fvy4xNkXgAEDF_D?format=jpg&name=medium https://pbs.twimg.com/media/Fvy4xNkXgAEDF_D?format=jpg&name=... If the extrapolation is not too flawed, it looks like PaLM 2-S might be about 120B, PaLM 2-M 180B, PaLM 2-L 280B. Still, I would expect GPT-4 trained for way longer than Chinchilla, so it could be smaller than even PaLM 2-S.
- MacsHeadroom 3y agoThey said the smallest PaLM 2 can run locally on a Pixel Smartphone. There's no way it's 120B parameters. It's probably not even 12B.
- espadrine 3y agoI am talking about the 3 larger models PaLM 2-S, PaLM 2-M, and PaLM 2-L described in the technical report. At I/O, I think they were referencing the scaling law experiments: there are four of them, just like the number of PaLM 2 codenames they cited at I/O (Gecko, Otter, Bison, and Unicorn). The largest of those smaller-scale models is 14.7B, which is too big for a phone too. The smallest is 1B, which can fit in 512MB of RAM with GPTQ4-style quantization. Either that, or Gecko is the smaller scaling experiment, and Otter is PaLM 2-S.
- MacsHeadroom 3y agoMy Pixel 6 Pro has 12GB of RAM and LLaMA-13B only uses 9GB in 4bit.