4 ms·
RedPajama ran on the Summit supercomputer (https://en.wikipedia.org/wiki/Summit_(supercomputer) https://en.wikipedia.org/wiki/Summit_(supercomputer)) NVIDIA V10
by emadm 3y ago
RedPajama ran on the Summit supercomputer (https://en.wikipedia.org/wiki/Summit_(supercomputer) https://en.wikipedia.org/wiki/Summit_(supercomputer)) NVIDIA V100s/PowerPC chips as part of the INCITE grant which necessitated variations from the LLaMA training parameters.
This led to differences in evals, now they are bringing in more modern chips.
Aside from some tokeniser differences this is a drop in replacement for existing LLaMA that matches performance.
- behohippy 3y agoHey emad, thanks for SD and this! What's the plan if Meta does Apache 2.0 for LLaMA? Just keep going and making the 30b and 65b or build different models?
- emadm 3y agoHad a nice chat with Yann last week, we will release complementary stuffs. I don't think 30b and 65b are useful given what we do, the key is optimising models for consumer & swarming them. As for SD.. Maybe try the bot on the discord server testing the new version: https://discord.com/invite/stablediffusion https://discord.com/invite/stablediffusion
- cypress66 3y ago30b 4bit is definitely useful because it runs on 3090s/4090s. You see a lot of people running 30b models. The jump in quality is very significant as well. 65b is definitely a lot less common.
- Vetch 3y agoThe tokenizer differences are major as LLMs are sensitive to whitespace handling. If I am reading the github page properly, OpenLLama failed to learn how to model code properly? Code contains many implicit reasoning tasks. What other differences are there? The page doesn't mention how numbers are handled. These are two major things that impact model reasoning and numeric ability.
- emadm 3y agoCode is main thing, it has some tradeoffs. It tunes well on code though and the code ai team at stability ai are working on stuff. We can now set and forget runs so will have a better dataset for the next 13b and different tokeniser, this was meant to match as close as possible to be a drop in.