3 ms·
Open AI Whisper was probably made to make transcripts of videos that will be used to train future AI models. The text in itself won't be that interesting, the
by macrolime 4y ago
Open AI Whisper was probably made to make transcripts of videos that will be used to train future AI models.
The text in itself won't be that interesting, the magic happens once you essentially train three different token predictors, one that predicts image tokens (16x16 pixels) and then combine that to predict video frames, one that predicts audio tokens and one that predicts text tokens. Then you use cross-attention between these predictors. To train this model you first pre-train the text predictor, after that's done you continue training the text predictor from the transcribed videos, while combining it with the video predictor and audio predictor with cross-attention.
Such a model will understand physics and actions in the real world much better than GPT-4, combined with all the knowledge from all the text on the internet it should turn out to be something quite interesting.
I think there probably doesn't exist enough compute yet to train such a model on something like all of YouTube, but I wouldn't be surprised if GPT-5 is a first step in this direction.
- candiodari 4y agoBut this can never work. Or it will just do one thing, and one thing only. All this does is summarize information produced by humans, to "extract value" from it. It cannot produce information itself. In other words: to make sure all the money goes to the plumber finding service, never to the plumber. These AIs serve to extract value from these summaries through "enshittification". https://pluralistic.net/2023/01/21/potemkin-ai/ https://pluralistic.net/2023/01/21/potemkin-ai/ It will massively increase the value of human interaction ... and the cost. It may even go so far that actual professionals start hiding from the internet, to hide from these models and avoid enshittification that way. We may have to start going back to a visiting a local cafe to find a plumber. I don't even think that is such a bad thing. Any AI that wants to actually learn more than just language (from summarization to bullshitting) will need to take actions in the real world and learn from them.
- HarHarVeryFunny 4y ago> Any AI that wants to actually learn more than just language (from summarization to bullshitting) will need to take actions in the real world and learn from them. Well, that's one additional source of knowledge, but we haven't even got into other sources like voice, video, simulation data, etc (let alone the interesting stuff like financial market data, aggregate human behavior, etc, etc). GPT-4 is only just starting to support image data, and who knows what GPT-5, 6.. will bring. Robotic embodiment is maybe better thought of as robotics, not an necessity for advancing AI in dozens of useful ways. But of course not every AI needs to be an expert in every domain. If you want to build a robot then it'll need to learn via interaction, but there's are tons of applications for even these 1st-gen language-only systems - look at all the applications that people have already found for them, and what is starting to be done with LangChain and plugins.