7 ms·
OLMo: Accelerating the Science of Language Models [pdf]
- gardenfelder 3y agohttps://huggingface.co/allenai/OLMo-7B https://huggingface.co/allenai/OLMo-7B Edit: add https://github.com/allenai/OLMo https://github.com/allenai/OLMo
- artninja1988 3y agoFeels like there must be 40 or so distinct open source llms now. What gives? We need some more new text to image models too... :(
- refulgentis 3y agoLanguages, sizes, and degrees of open-ness.
- chuckhend 3y agoThere's some more commentary on their open-ness in this blog too https://www.interconnects.ai/p/olmo https://www.interconnects.ai/p/olmo
- dwagnerkc 3y agoThat post also very helpfully links to another paper they published alongside the OLMo paper just on the dataset. Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research https://arxiv.org/abs/2402.00159 https://arxiv.org/abs/2402.00159
- chuckhend 3y agoThe training datasets are also available, which sets them apart a bit IMO. https://huggingface.co/datasets/allenai/dolma https://huggingface.co/datasets/allenai/dolma
- hedgehog 3y agoThere are few that are >1B params, competitive, and "open source" in the sense that the necessary ingredients to re-train are available. Models like Llama and thus its descendants (including Mistral's public models) have weights available but not the training data.
- wokwokwok 3y agoIf you read around, training a 7B model costs on the order of $85,000; the 1.4 stable diffusion release cost around $600,000 to train. You don't see a lot of 70B or larger models being released for the same reason; it's expensive. We should just be grateful for what we're getting right now: basically, people are spending 100s of thousands of dollars on training and giving the results away for free. Hugging face is hosting them for free. ollama is hosting them for free. People are writing free inference engines (eg. llama.cpp) and giving them away. Don't complain. We've got it pretty damn good right now.
- alchemist1e9 3y ago> If you read around, training a 7B model costs on the order of $85,000; the 1.4 stable diffusion release cost around $600,000 to train. That seems remarkably cheap actually and likely getting cheaper fairly quickly with improvements in training efficiencies I’d imagine.
- sjwhevvvvvsj 3y agoOn the other hand, the systems are trained on “free” data so it kinda should be public property by default. Claiming it’s fair use to suck up the entire web and pay wall the derived result is absurd argument. We all created the lifeblood of LLM and we’re entitled to the product.
- visarga 3y ago> We all created the lifeblood of LLM and we’re entitled to the product. sounds so nice, yet there are going to be objections, NYT for example doesn't think we all should be entitled to the product
- sjwhevvvvvsj 3y agoOf course, that is partially my point: if OpenAI et al wants to make the argument anything online is fair game, then they should release the weights. If not, they have no leg to stand on.
- 3y ago
- thawab 3y agoOpen source means i have documentation to reproduce the same results. This is only true with tinyllama and this model. The other models (llama, mistral) are free to use and not open source.
- arugulum 3y agoThe Pythia models have all the training data, code, and configurations available.
- jerrygenser 3y agoPretty cool that it runs on and and Nvidia
- shwaj 3y agoNot sure if you’re downvoted for the typo: “and” instead of “AMD”?
- jerrygenser 3y agoYes I meant AMD.
- alchemist1e9 3y agoThey detail the energy used and therefore estimated carbon emissions which is interesting. When I estimate the raw electricity cost using 7-20 cents per kWh for US commercial rates, then we are only talking about $16-50k for electricity, that seems pretty small! Is my math wrong? Is there any information on how much the computing costs were for renting the clusters? Is the barrier to entry for a 7B model only a couple $100K? EDIT: https://news.ycombinator.com/item?id=39223467#39224534 https://news.ycombinator.com/item?id=39223467#39224534 Perhaps only $85K total
- anonylizard 3y agoDespite the typical complaints about "X new thing harming the environment!!!", LLMs are as friendly as it gets, it 1. Consumes a minor amount of electricity (Data centers is only 2% of US electricity use, and currently AI is maybe only 5-10% of that). Its trivial compared to say metal smelting. 2. Consume water for cooling. That's it, there is 0 direct pollution generated from AI, and even the water use is very minor compared to say farming, and can be improved via more water efficient cooling techs. The main concern is the scaling speed. As LLMs scale up 10x, 100x, 1000x, those previously very minor electricity costs can quickly become grid impacting in a decade.
- marmaduke 3y agoI can't buy this kind of argument anymore. How about the external effect of AI steering the entire semiconductor industry to increase GPU/NPU capacity?
- vortegne 3y agoThis kind of argument is actually totally valid. But only if you subscribe to the current meta of widely accepted handwaving. Externalities are never a part of capitalist math. Non-trivial consequences can never hurt if one never looks further than their own nose.
- StopTheTechies 3y ago> we are only talking about $16-50k for electricity, that seems pretty small I suppose this depends greatly on how you view the utility of LLMs. In a capitalist sense, sure—there's great utility here persuading VCs to part with their coins and jobs to be replaced with correspondingly larger profit margins. But the opportunity cost of not solving major problems most of humanity can agree on seems nearly incalculably large. Not that capitalists give a shit.
- bravura 3y ago"We intend to follow up on this release with another one soon that includes the following: ... Weights & Biases logs for our training runs." That's amazing. I've never seen that before in a paper of this quality. Or, any paper at all.
- marvinalone 3y agoWeights & Biases for OLMo 7B are now out: https://wandb.ai/ai2-llm/OLMo-7B/reports/OLMo-7B--Vmlldzo2NzQyMzk5 https://wandb.ai/ai2-llm/OLMo-7B/reports/OLMo-7B--Vmlldzo2Nz...
- swyx 3y agoi think huggingface and facebook have both offered this level of detail in the past? still great though
- arugulum 3y agoEleutherAI as well.
- gillesjacobs 3y agoIt's more common than you think. I did the same for one of my research papers.
- nl 3y agoIt's very interesting that they went to the effort of doing complete end-to-end runs on both NVidia and AMD hardware. A pity they didn't release the speed of training, but the software is now there for someone else (not under benchmark embargo) to do that.
- nl 3y agoWho will be the first to do a useful Instruct-trained variant? It's a pity the Mistral 7B Instruct 0.2 dataset isn't available because I've found that a much higher quality than any of the finetunes around, and I suspect we'll have to rely on the same groups doing finetunes for this.
- bugglebeetle 3y agoNous just released their full instruction tuning dataset, so I dunno why someone with enough compute couldn’t do this.
- cosmojg 3y agoAnd Capybara be lookin' fiiine for tuning too. Seriously, though, you're right. These are some of the highest quality generative datasets in existence, and I'm surprised more isn't being done with them.
- nl 3y agoThe Nous finetunes of Mistral benchmark well but in practice seem worse than the original Mistral versions IMHO. Of course we don't know how to measure this so respect to them for the benchmark performance.
- casercaramel144 3y agoI'm sorry, I don't understand the exact contribution here? There's many tutorials on how to train a language model. If it's a repository of SOTA techniques for training, this will be outdated in at max 3 months, and anyways the ground shifts under you in this field so you might as well read Arxiv all day if your intention is to keep up with SOTA.
- tkellogg 3y agoresearchers don't read tutorials, they cross check each other's work. You need details to do that.
- casercaramel144 3y agowdym by cross check each others work? Surely just reporting the final loss is good enough if that's the intention. The final end goal is lower loss anyways so it's not even a bad metric.
- chuckhend 3y agoIt looks like this team gave us everything we need to reproduce their models, the actual artifacts needed to reproduce it. As far as I can tell, they share the data and every step along the way to final model...not just describing what they did.