4 ms·
I think part of the problem for LLMs is that they don't operate in an easily scoreable "game". We can make them work well for optimizing next token(s) log likel
by dbmikus 3y ago
I think part of the problem for LLMs is that they don't operate in an easily scoreable "game". We can make them work well for optimizing next token(s) log likelihood, but how do we judge quality of the output for the task it was made for?
Then after that, there are other challenges, such as do LLMs have a world model that lets them think how to attain a reward?
Wherever you can have an automated feedback mechanism, you can start to address the first problem and allow for the AI to explore more of the "tree". These are situations like coding (you can run the code and evaluate the output), or situations where you can let the LLM crowd-source user feedback.
For models that have some actual world model, who can reason across modalities, and who can plan, LeCunn talks about this often. The videos/slides here[1] were good content.
[1]: https://www.ece.uw.edu/news-events/lytle-lecture-series/ https://www.ece.uw.edu/news-events/lytle-lecture-series/
- visarga 3y agoYes, whatever is not written in any books, AI will have to learn directly from the feedback generated by the environment to its actions. AlphaZero discovered go from scratch, and beat us who invented the game and had thousands of years to practice it. This is how powerful a teacher can be the environment.
- Zambyte 3y ago> I think part of the problem for LLMs is that they don't operate in an easily scoreable "game". Unfortunately "social media" is the gamified environment for language.
- Jensson 3y agoBut it has human voters, you can't train a model using human voters to vote during iterations, and AI aren't good enough to replace human voters. Not sure if AI voters even can result in a model smarter than the voting model.
- Zambyte 3y ago> you can't train a model using human voters to vote during iterations Why not? Just comment multiple times on a post in different ways. Score outputs with more of a desired response higher than outputs with less desired responses. Scale that up site-wide on multiple sites, and it seems like you have a pretty powerful way to get human feedback...
- Jensson 3y agoThese models trains on billions of examples, having bots posting billions of posts on different social media sites every time you train a new model probably isn't a viable strategy. You would get banned real quick since most of those will be really low quality in the early stages.
- vineyardmike 3y agoSo I agree with your overall sentiment, BUT we already have good LLMs, so I think we’ve crossed the bridge from “bad low quality responses in the beginning” and I also suspect there’s more Social Media bots than we’d like to admit.
- Zambyte 3y agoI think we're simply talking about different things. Obviously you wouldn't want to start with a model of pure noise when interacting with real humans like that. I am describing using RLHF to fine tune existing models. That is a way to gamify training.
- user_7832 3y agoI personally feel like these problems can be broken into two types - one where the output is expected/deterministic, and one where creativity is a virtue. Asking Siri what the current temperature is (only 1 correct answer) is an example of the former, but asking chatGPT to write an email is closer to the latter. Tree type algorithms are better when you don't have (millions of english words)^10 for a 10 word response, and where there are multiple "correct" answers.
- dontupvoteme 3y agoI mean you can make any set of LLMs produce any set of output you want. The problem is that it isn't terribly efficient and you have to filter between the real stuff and the hallucinations.
- roenxi 3y agoQualitatively, the answer seems to be to train for novelty. Which happens to also be how humans accomplish much of the complex stuff. Eg, One interpretation of art is we have 8 billion or so humans with a finely trained neural net for recognising stuff. The artist is looking for novel ways of triggering those nets, and if they find something inspired then it is art. A great artist generally isn't trying to reproduce something that is known, they're trying to explore novel areas of a medium. So the real trick to building the world model is coming up with a good novelty metric. Still hard, but easier than developing a reward function. That gets the part of the training done that establishes a world model, then I'd assume it is possible to train that model to do a task by rewarding specific outcomes that it already knows how to achieve.
- NBJack 3y agoIf it's a game, I think the old phrase "everything is made up and the points don't matter" applies well here. It's really hard to score a phrase for truthfulness. Even among humans often it's ambiguous; we have courts of law in many societies to try and fix some of that. I admit I am disappointed that we don't see a similar amount of work these days in what was once dubbed 'expert systems', though perhaps the results just aren't as flashy.