7 ms·
Once you've trained on the internet and most published books (and more...) what else is there to do? You can't scale up massively anymore.
by cleandreams 3y ago
Once you've trained on the internet and most published books (and more...) what else is there to do? You can't scale up massively anymore.
- deleted 3y ago[deleted]
- m4jor 3y agoThey didn't train it on the entire internet tho, only a small amount (in comparison to entire internet). Still plenty they could do.
- JasonZ2 3y agoVideo. > YouTubers upload about 720,000 hours of fresh video content per day. Over 500 hours of video were uploaded to YouTube per minute in 2020, which equals 30,000 new video uploads per hour. Between 2014 and 2020, the number of video hours uploaded grew by about 40%.
- sottol 3y agoBut what are you mostly "teaching" the LLM then? Mundane everyday stuff? I guess that would make them better at "being average human" but is that what we want? It already seems that prompting the LLM to be above-average ("pretend to be an expert") improves performance.
- dougmwne 3y agoThis whole conversation about training set size is bizarre. No one ever asks what’s in the training set. Why would a trillion tokens of mundane gossip improve a LLMs ability to do anything valuable at all? If a scrape of the general internet, scientific papers and books isn’t enough, a trillion trillion trillion text messages to mom aren’t going to change matters.
- Animats 3y agoRight. They've already sucked in most of the good general sources of information. Adding vast amounts of low-quality content probably won't help much and might degrade the quality of the trained model.
- rvnx 3y agoVideo content (I don't know why someone flagged Jason for saying such, he is totally right)
- bheadmaster 3y agoLooking at his post history, seems like he was shadowbanned.
- machdiamonds 3y agoIlya Sutskever (OpenAI Chief Scientist): "Yeah, I would say the data situation is still quite good. There's still lots to go" - https://youtu.be/Yf1o0TQzry8?t=685 https://youtu.be/Yf1o0TQzry8?t=685 There was a rumor that they were going to use Whisper to transcribe YouTube videos and use that for training. Since it's multimodal, incorporating video frames alongside the transcriptions could significantly enhance its performance.
- it_citizen 3y agoI am curious how much video-to-text content represent compared to pure text. I have no idea.
- deleted 3y ago[deleted]
- neel8986 3y agoAnd why will google allow them to do that at scale?
- throwaway5959 3y agoWhy would they ask Google for permission?
- neel8986 3y agoBecause youtube is owned by google and google can stop download at scale
- HDThoreaun 3y agoCan google stop them? It’s trivial to download YouTube videos
- unionpivo 3y agoIt’s trivial to download some YouTube videos. But I am quite sure that if you start doing it at scale, google will notice. You could be sneaky, but people in this business talk (since they know another good paying job is just around the corner) so It would likely come out.
- neel8986 3y agoYoutube. This is where Google have huge advantage having largest collection of user generated video
- sebzim4500 3y agoYeah, but it's not like the videos are private. Surely Amazon has the real advantage, given they have a ton of high quality tokens in the form of their kindle library and can make it difficult for OpenAI to read them all.
- mrtksn 3y agoYou can transcribe all spoken words everywhere and keep the model up to date? Keep indexing new data from chat messages, news articles, new academic work etc. The data is not finite.
- spaceman_2020 3y agoWhat about all the siloed content kept inside corporate servers? You won't get normal GPT to train on it, of course, but IBM could build a "IBM-bot" that has all the GPT-4 dataset + all of IBM's internal data. That model might be very well tuned to solve IBM's internal problems.
- treis 3y agoI don't think you can just feed it data. You've got to curate it, feed it to the LLM, and then manually check/further train the output. I also question that most companies have the volume and quality of data worth training on. It's littered with cancelled projects, old products, and otherwise obsolete data. That's going to make your LLM hallucinate/give wrong answers. Especially for regulated and otherwise legally encumbered industries. Like can you deploy a chat bot that's wrong 1% or 0.1% of the time?
- spaceman_2020 3y agoWell, IBM has 350k employees. If training a LLM on curated data costs tens of millions of dollars but ends up reducing headcount by 50k, it would be a massive win for any CEO. You have to understand that all the incentives are perfectly aligned for corporations to put this to work, even spending tens of millions in getting it right. The first corporate CEO who announces that his company used AI to reduce employee costs while increasing profits is going to get such a fat bonus that everyone will follow along.
- Vrondi 3y agoSince Chat-GPT-4 is being integrated into the MS Office suite, this is an "in" to corporate silos. The MS cloud apps can see inside a great many of those silos.
- spaceman_2020 3y agoIf you were devious enough, you could be listening in on billions of phone conversations and messages and adding that to your data set. This also makes me doubt that NSA hasn't already cracked this problem. Or that China won't eventually beat current western models since it will likely have way more data collected from its citizenry.
- PUSH_AX 3y agoI wonder what percentage of phone calls would add anything meaningful to models, I imagine that the nature of most phone calls are both highly personal and fairly boring.
- midland_trucker 3y agoThat's a fair point. Not at all like training on Wikipedia in which nearly every sentence has novelty to it. Then again it would give you data on every accent in the country, so the holy grail for modelling human speech.
- kolinko 3y agoYou can generate textual examples that teach logic, multi-dimensional understanding and so on. Similar to the ones that are in math books, but in a massive scale.
- nabnob 3y agoReal answer? Buy proprietary data from social media companies, credit card companies, retail companies and train the model on that data.
- eukara 3y agoCan't wait for us to be able to query GPT for peoples credit card info
- sebzim4500 3y agoI doubt they have trained on 0.1% of the tokens that are 'easily' available (that is, available with licencing deals that are affordable to OpenAI/MSFT). They might have trained on a lot of the 'high quality' tokens, however.
- fpgaminer 3y ago> Once you've trained on the internet and most published books (and more...) what else is there to do? You can't scale up massively anymore. Dataset size is not relevant to predicting the loss threshold of LLMs. You can keep pushing loss down by using the same sized dataset, but increasingly larger models. Or augment the dataset using RLHF, which provides an "infinite" dataset to train LLMs on. Limited by the capabilities of the scoring model which, of course, you can scale the scoring model infinitely so again the limit isn't dataset size but training compute.
- midland_trucker 3y ago> Dataset size is not relevant to predicting the loss threshold of LLMs. You can keep pushing loss down by using the same sized dataset, but increasingly larger models. Deepmind and others would disagree with you! No-one really knows in actual fact. [1] https://www.deepmind.com/publications/an-empirical-analysis-of-compute-optimal-large-language-model-training https://www.deepmind.com/publications/an-empirical-analysis-...
- fpgaminer 3y agoI don't recall the Chinchilla paper disputing my point. They establish "training-compute optimal" scaling laws, but none of their findings suggest that loss hits any kind of asymptote.
- midland_trucker 3y agoPerhaps we're talking past each other, is "loss threshold" a specific term in LLM literature? Merely pointing out that the debate as to whether we are compute or data limited (OP) has not concluded at all; There are lots of compelling theories on relationship between the two.
- narrator 3y agoYou could have it start talking to itself in the way that AlphaGO learns to get better at Go. All that needs to be done is find some fitness function that indicates that useful knowledge has been produced. In Go and Chess this is easy. It can start posting synthesized ideas on social media and see how many likes it gets. Coupled with a metric containing dissimilarity to current information, this could be a useful way to progress to superhuman insights.
- lumenwrites 3y agoVideos - all of youtube, all the movies, everything that's ever been captured on film. Transcribe the audio, automatically describe the images and try to predict the next one.
- ChatGTP 3y agoThe problem is, movies are completely full of misinformation and inaccuracies. Maybe if you trained it on movies before CGI existed ?
- kvetching 3y agopeople seem to have forgotten about the multi-modal GPT-4 There's a ton of potential left on the table. The question is if transformers have hit their limit with GPT-4 or not. It's a pretty simple equation when you think about it this way and why Sam would say they have hit their limit. Sam is basically Microsoft and they want to retain their lead. Once Google learns to put their data to use correctly, it's almost guaranteed game over for OpenAI if they want it to be.