3 ms·
I'm not sure if I'm missing something from the paper, but are multi-billion parameter models getting called "small" language models now? And when did this parad
by yujian 3y ago
I'm not sure if I'm missing something from the paper, but are multi-billion parameter models getting called "small" language models now? And when did this paradigm shift happen?
- Chabsff 3y agoNowadays, small essentially means realistically useable on prosumer hardware.
- nathanfig 3y agoRelative term. In the world of LLMs, 7b is small.
- hmottestad 3y agoAll the llama models, including the 70B one can run on consumer hardware. You might be able to fit GPT-3 (175B) at Q4 or Q3 on a Mac Studio, but that's probably the limit for consumer hardware. At 4-bit a 7B model requires some 4GB of ram, so that should probably be possible to run on a phone, just not very fast.
- sa-code 3y agoGpt 3.5 turbo is 20B
- kristianp 3y agoI doubt that. What's your source?
- sa-code 3y agoThere was a paper published by Microsoft that seemed to leak this detail. I'm on mobile right now and don't have a link but it should be searchable
- nl 3y agoThe paper was https://arxiv.org/abs/2310.17680 https://arxiv.org/abs/2310.17680 It has been withdrawn with this note: > Contains inappropriately sourced conjecture of OpenAI's ChatGPT parameter count from this http URL, a citation which was omitted. The authors do not have direct knowledge or verification of this information, and relied solely on this article, which may lead to public confusion (the noted URL is a just a Forbes blogger with no special qualifications that would make what he claimed particularly credible).
- moffkalast 3y agoWhen 175B, 300B, 1.8T models are considered large, 7B is considered small.