8 ms·
It's fantastic that more orgs are releasing open-source models trained on more than 300B or so tokens. Here's my take from the details I could find. Pros -
by vikp 3y ago
It's fantastic that more orgs are releasing open-source models trained on more than 300B or so tokens. Here's my take from the details I could find.
Pros
- 4096 context width (vs 2048 for llama, gpt-j, etc)
- 3B to 65B released or in progress
- RL tuned models available
- Trained on more tokens than existing non-llama models
- 128 head dim, so can use flash attention (unlike GPT-J)
Cons
- No benchmarks released, or details about the model
- Somewhat restrictive license on the base models, and NC license on the RL models
- Small models only trained on 800B tokens, compared to 1T for llama-7B, and potentially more for other upcoming alternatives (RedPajama, etc). I'd like to see their loss curves to see why they chose 800B.
High-level, this is likely to be more accurate than existing non-llama open source models. It's hard to say without benchmarks (but benchmarks have been gamed by training on benchmark data, so really it's just hard to say).
Some upcoming models in the next few weeks may be more accurate than this, and have less restrictive licenses. But this is a really good option nonetheless.
- capableweb 3y ago> - 3B to 65B released or in progress Seems they want to do 3B to 175B, although 175B is not in progress yet.
- ipsum2 3y agoIt's not efficient to do 175B. Training a smaller model (65B) on more data gives better performance for the same compute.
- tempaccount420 3y agoIf you want it to just regurgitate training data, sure. But more parameters will always be better for more complex tasks.
- thewataccount 3y ago> But more parameters will always be better for more complex tasks. I think you should checkout this paper which discusses the relationship of performance and the ratio of training tokens to parameter count. https://arxiv.org/abs/2203.15556 https://arxiv.org/abs/2203.15556
- tempaccount420 3y agoStableLM already has an optimal parameter number to tokens ratio, so what's your point? They should train the 65B model on even more tokens? > StableLM is trained on a new experimental dataset built on The Pile, but three times larger with 1.5 trillion tokens of content
- thewataccount 3y agoIf I understand correctly, based on their prediction in Table 3 on page 8, they do have enough tokens, but they also need over a magnitude more compute time. > It's not efficient to do 175B. Training a smaller model (65B) on more data gives better performance for the same compute. This is OP's comment you replied to - so I was responding under OP's context that the amount of compute time would be the same, which I apologize I didn't make clear, and my response was very poorly worded. My intent was to link the paper because I think it supports OP's statement that for the same amount of compute time and a token ratio, the performance of a smaller model will be better then a larger one (assuming they haven't converged yet which they haven't at this size). > If you want it to just regurgitate training data, sure. This paper was about showing Chinchilla performing with models many times larger then itself, showing you don't need to have a 175B size model for more performance then "regurgitating training data"
- wokwokwok 3y ago> you don't need to have a 175B size model… Sure, that’s true. …but, a fully trained larger model is going to be better. There only reasonable reason to prefer a smaller model is because it’s cheaper and less intensive to train. It sounds a lot like you’re saying “small models are just as good” … which is false. No one believes that. For a given compute budget an under trained large model and a well trained small mode may be comparable, right? …but surely, the laws of diminishing returns applies here? There’s an upper bound to how good your smaller model can ever be, right? Over time, someone can take a larger model which is under trained and refine that model right? The “small model is just as good” narrative only holds up for a fixed once only training of a model for a fixed compute budget at the moment of release. Over all of time that compute budget is not fixed.
- sebzim4500 3y agoDepends on your compute budget.
- kiraaa 3y agoand also easy to deploy
- GaggiX 3y ago>Small models only trained on 800B tokens "These models will be trained on up to 1.5 trillion tokens." on the Github repo. https://github.com/stability-AI/stableLM/#stablelm-alpha https://github.com/stability-AI/stableLM/#stablelm-alpha
- Taek 3y agoDevs confirmed that the small ones use 800B, 1.5T is for the large ones
- GaggiX 3y ago@thunderbird120 asked a Stability employee and say that the plan is going to keep training the models up to 1.5T. So I don't know where do you read this.
- nickthegreek 3y agohttps://github.com/Stability-AI/StableLM#stablelm-alpha https://github.com/Stability-AI/StableLM#stablelm-alpha shows that the 3b and 7B had 800b training tokens.
- Taek 3y agoThat may be, but the weights you can download today were trained on 800B
- GaggiX 3y agoyes of course that's why they use "will be trained" on the GH repo.
- sroussey 3y agoI think they are “checkpoint” models in this case. Will be fun to compare when completed!
- oehtXRwMkIs 3y ago
- HarHarVeryFunny 3y agoThey mention 1.5T training tokens, perhaps for the largest model only ?
- vikp 3y agoIt's unclear which models will be trained to 1.5T tokens. The details of how many tokens each model saw in training are on Github - https://github.com/stability-AI/stableLM/ https://github.com/stability-AI/stableLM/ . But only for the ones that have been released.
- thunderbird120 3y agoI just asked a stability employee and they said the the current models ran into an overfitting issue probably due to some duplicated data somewhere in their dataset, which consists of 1.5T tokens. The 800B tokens is the number of tokens they've been trained on so far. The plan is to keep going and train on the rest of the data once the issue is resolved.
- HarHarVeryFunny 3y agoI've asked this question in a few places, and never been able to get an answer, maybe you know... Q: Why are these LLMs trained on a single epoch, and perform worse if the dataset is repeated ? This seems maybe related to suspecting data duplication as a cause of overfitting. Why don't LLMs need multi-epoch training at a low learning rate to generalize? If they are managing to learn from a single epoch, that sounds more like they may be memorizing!
- thunderbird120 3y agoNever repeating your training data is what you'd ideally like to do for training basically any ML model. If you do that you don't really need to worry about overfitting since the model is constantly trying to fit a stream of new data. To reduce its training error it actually has to model the structure of the data rather than just memorizing it since each training step will involve data it has never seen before. Larger models are more prone to overfitting but also learn several orders of magnitude faster. If you can use larger models without being concerned about overfitting it's generally desirable to do so. It's just that most tasks don't actually have enough data to support doing that. Thankfully, text modeling does have enough data.
- whimsicalism 3y ago> Small models only trained on 800B tokens, compared to 1T for llama-7B LLaMA is trained far beyond chinchilla optimality, so this is not as surprising to me.
- anentropic 3y agoAccording to this LLaMA still didn't go far enough: https://www.harmdevries.com/post/model-size-vs-compute-overhead/ https://www.harmdevries.com/post/model-size-vs-compute-overh...
- whimsicalism 3y agoYep, it depends on what your goal is.
- cubefox 3y agoThis doesn't say that LLaMA didn't go far enough.
- anentropic 3y agoNot exactly, but it did say they could have gone further than they did without wasting time and energy on infinitesimally small gains though
- dragonwriter 3y agoBut Chinchilla optimality, while an interesting result, is a strange target for most practical purposes. Training is one time, inference is many times; not training past the point where its cheaper to training a larger model for the same (proxy for) quality discounts to zero the import of the cost of inference.
- whimsicalism 3y agoYep, but if stability has the goal of training the best possible model then that would explain the choices they made.
- sebzim4500 3y ago>- No [...] details about the model You can see the model architecture here https://github.com/Stability-AI/StableLM/blob/main/configs/stablelm-base-alpha-7b.yaml https://github.com/Stability-AI/StableLM/blob/main/configs/s...
- beecafe 3y ago[dead]
- DustinBrett 3y agoI'm wondering what the sweet spot for parameters will be. Right now it feels like the Mhz race we had back in the CPU days, but 20 years later I am still using a 2-3GHz CPU.
- lhl 3y agoI think "sweet spot" is going to depend on your task, but here's a good recent paper that may give you some more context on thinking about training and model sizes: https://www.harmdevries.com/post/model-size-vs-compute-overhead/ https://www.harmdevries.com/post/model-size-vs-compute-overh... There have also been quite a few developments on sparsity lately. Here's a technique SparseGPT which suggests that you can prune 50% of parameters with almost no loss in performance for example: https://arxiv.org/abs/2301.00774 https://arxiv.org/abs/2301.00774
- version_five 3y agoI was wondering if the longer training thing was a similar phenomenon to the double-descent we see in other deep learning models. Training for a really long time can improve generalization (as can adding more parameters) - but I don't know enough about LLM architecture to know if that's relevant here. My skim of the blog post led me to think it's proposing a different mechanism (scaling laws).
- Taek 3y agoWell, based on all the data we have available now it seems like you don't get much benefit yet from going above 200 billion.
- lhl 3y agoFYI, I'm running lm-eval now w/ the tests Bellard uses (lambada_standard, hellaswag, winogrande, piqa,coqa) on the biggest 7B an 40GB A100 atm (non-quantized version, requires 31.4GB) so will be directly comparable to what various LLaMAs look like: https://bellard.org/ts_server/ https://bellard.org/ts_server/ (UPDATE: run took 1:36 to complete run, but failed at the end with a TypeError, so will need to poke and rerun). I'll place results in my spreadsheet (which also has my text-davinci-003 results): https://docs.google.com/spreadsheets/d/1kT4or6b0Fedd-W_jMwYpb63e1ZR3aePczz3zlbJW-Y4/edit#gid=0 https://docs.google.com/spreadsheets/d/1kT4or6b0Fedd-W_jMwYp...
- lunixbochs 3y agoAre you using https://github.com/EleutherAI/lm-evaluation-harness https://github.com/EleutherAI/lm-evaluation-harness?
- lhl 3y agoYeah, although looks like it currently has some issues with coqa: https://github.com/EleutherAI/lm-evaluation-harness/issues/238 https://github.com/EleutherAI/lm-evaluation-harness/issues/2... There's also the bigscience fork, but I ran into even more problems (although I didn't try too hard) https://github.com/bigscience-workshop/lm-evaluation-harness https://github.com/bigscience-workshop/lm-evaluation-harness And there's https://github.com/EleutherAI/lm-eval2/ https://github.com/EleutherAI/lm-eval2/ (not sure if it's just starting over w/ a new repo or what?) but it has limited tests available
- guywithabowtie 3y agoDo you also have results of GPT4 somewhere? or text-davinci-003-turbo
- lhl 3y agoI'm still on the waitlist for GPT-4 API access. Note, that text-davinci-003 cost about $90 to benchmark at $0.02/1K tokens, so if you're able to use a GPT-4 model (for completion and not just instruction) that'll probably be $270-$540 in credits to benchmark...
- swyx 3y ago> 128 head dim, so can use flash attention (unlike GPT-J) mind explaining why this is so attractive/what the hurdle is for the laypeople in the audience? (me)
- GaggiX 3y agoStandard attention has memory quadratic in sequence length, whereas FlashAttention has memory linear in sequence length. Also FalshAttention is faster.
- sroussey 3y agoSo there must be a downside to FlashAttention. What is it?
- lhl 3y agohttps://arxiv.org/abs/2205.14135 https://arxiv.org/abs/2205.14135 - Section 5 suggests that the biggest limitation is that custom CUDA kernels need to be coded on a per-GPU architecture basis.
- fpgaminer 3y agoFlashAttention is mathematically identical to standard attention, so in theory there's no downside. In practice, numerical inaccuracies of floating point mean that the results differ slightly. I don't know of any papers going in depth to analyze what impact those variances have in a range of real models, but generally speaking deep models handle slightly variances well. I've not noticed any difference in my applications training models. And tons of people use FlashAttention as a drop-in replacement on models trained on standard attention (e.g. using xformers in StableDiffusion). Also in practice FlashAttention is still relatively new so it isn't well supported in libraries yet. Until PyTorch 2.0 you had to either implement it yourself, or use something like xformers which comes with a bag of caveats. PyTorch 2.0 now has it built-in, and it's easy to use, but the implementation is incomplete so you can't, for example, use it with an attention mask (which is needed in LLMs, for example). tl;dr: Basically none, but it just isn't well supported yet.
- burtonator 3y agoWere you able to figure out if the RL models are going to be jailed? A 65B parameter model could be a bit frightening. That's 1/3rd the size of GPT3.
- sebzim4500 3y agoI'm sure there will be a bunch of different RL tuned versions of them, RLHF isn't that expensive. IIRC Microsoft has software that will do it for a few thousand dollars for a model that size. I'm sure someone will release a non-lobotomized version, maybe OpenAssistant.
- kiraaa 3y agoits not alway about the size, but yeah its really good!