3 ms·
Thank you for this amazingly insightful comment. A question as an ignorant layperson, if I may: They could have said it cost the same amount of carbon as 25
by warent 4y ago
Thank you for this amazingly insightful comment.
A question as an ignorant layperson, if I may:
They could have said it cost the same amount of carbon as 25 US drivers emit each year. In the grand scheme of things, that's nothing. That's one long-haul flight per trained model. And you just need to train them once. Querying the models cost far less.
Don't they need to continuously re-train these models as new information comes in? For example, how does Bing bot get new information? It seems like they would need to routinely keep it up-to-date with its own index.
- sdiacom 4y agoGoing by the screenshots in the linked tweets, it seems like it performs searches on Bing in order to obtain up-to-date information to answer its questions with, so there's probably not a need to re-train it daily. So the main question here might be "how much energy does it cost to keep a search engine up-to-date", which may not be cheap, either. There is probably a need to refresh it periodically to account for what the MMAcevedo fictional story [1] calls "context drift" -- the relevant search terms to infer from the query are themselves contextual. Say, if I ask Bing today "is Trump running for president", the right search term today could be "donald trump 2024 election", but ten years from now it might be "eric trump 2036 election". [1]: https://qntm.org/mmacevedo https://qntm.org/mmacevedo
- humanistbot 4y ago> Don't they need to continuously re-train these models as new information comes in? For example, how does Bing bot get new information? It seems like they would need to routinely keep it up-to-date with its own index. Sure, and thanks! Some keywords to look up are transfer learning, zero-shot learning, and fine tuning. These approaches focus on exactly this problem: not having to retrain the entire model from scratch to add new information. GPT-3's training data is 100 billion tokens of text, but to extend it by another 1 billion tokens of text is far closer to 1/100 the original cost. It actually wasn't the energy/carbon cost that motivated early work in this, it was more about adapting to new domains and letting people customize models for specific purposes. Image processing really adopted it first to great success. Orgs with resources trained really big models on all of ImageNet that needed server farms of GPUs, but they released it so that other people can use a single commodity GPU to fine-tune it for whatever their specific image processing task. Edit: now you can pay "Open"AI to fine tune their models for you, but only Microsoft has access to the raw model itself
- YeGoblynQueenne 4y ago>> Some keywords to look up are transfer learning, zero-shot learning, and fine tuning. These approaches focus on exactly this problem: not having to retrain the entire model from scratch to add new information. GPT-3's training data is 100 billion tokens of text, but to extend it by another 1 billion tokens of text is far closer to 1/100 the original cost. Well, if you fine-tune GTP-3 on another billion tokens you get a fine-tuned version of GPT-3, that's perhaps better at modelling those billion tokens. If you want a better GPT-3 you have to pre-train a new model, probably with a few more billion parameters. So transfer learning is not going to save the day here.
- YeGoblynQueenne 4y agoYeah, those models need to be retrained often. Not to update an index, but to keep up to date with the new content on the internet, and of course to create new, improved models with more parameters: note that after GPT, we had GPT-2, GPT-3, GPT-3.5 (a.k.a. ChatGPT) and rumours are we're now at GPT-4 (powering Bing search). Meanwhile, each of those models was trained in multiple versions, with different numbers of parameters; for example, the GPT-3 that first broke the barrier of hype was the largest, at 175 billion parameters, of four or five models. Plus, it seems reasonable that OpenAI, Google et al. are retraining models every once in a while to correct mistakes (e.g. OpenAI say in their GPT-3 paper that they couldn't retrain their 175Gp model to correct a mistake, because of the high cost of training, but that was a couple of years ago and they have clearly trained multiple large models since, so why not GPT-3 again? Except of course they're not very -cough- open, about those things, so we can't know for sure). Bottom line, the cost of training "GPT-3" varies a lot and is paid multiple times. ________ GPT-3 paper for ref: https://arxiv.org/abs/2005.14165 https://arxiv.org/abs/2005.14165