4 ms·
AlphaZero demonstrates that more human-generated data isn't the only thing that makes an AI smarter. It uses zero human data to learn to play Go, and just itera
by robbrown451 3y ago
AlphaZero demonstrates that more human-generated data isn't the only thing that makes an AI smarter. It uses zero human data to learn to play Go, and just iterates. As long as it has a way of scoring itself objectively (which it obviously does with a game like Go), it can keep improving with literally no ceiling to how much it can improve.
Pretty soon ChatGPT will be able to do a lot of training by iterating on its own output, such as by writing code and analyzing the output (including using vision systems).
Here's an interesting thing I noticed last night. I have been making a lot of images that have piano keyboards in them. DALL-E 3 makes some excellent images otherwise (faces and hands mostly look great), but it always messes up the keyboards, as it doesn't seem to get how black keys are in alternating groups of two and three.
But I tried getting chatgpt to analyze an image, using its new "vision" capabilities, and the first thing it noticed was that the piano keys were not properly clustered. I said nothing about that, I just asked it "what is wrong with this image" and it immediately found that. What if it could feed this sort of thing back in, using similar logic to Alpha Zero?
That's just a tiny hint of what is to come. Sure, it typically needs human generated data for most things. It's already got thousands of times more than any human has looked at. It will also be able to learn from human feedback, for instance a human could tell it what it got wrong in a response (whether regular text, code, or image), and explain in natural language where it deviated from what was expected. It can learn which humans are reliable, so it can minimize the number of paid employees doing RLHF, using them mostly to rate (unpaid) humans who choose to provide feedback. Even if most users opt out of giving this sort of feedback, there will be plenty to give it new, good information.
- realistic2020 3y agoWith Alpha Go, you have a clear objective -- to win a game. How does that work for creative outputs?
- ToValueFunfetti 3y agoThe same way we do it. Verifying that an output is good is far easier than producing a good output. We can write a first draft, see what's wrong with it, make changes, and iterate on that until it's a final draft. And along the way we get better at writing first drafts.
- Jerrrry 3y ago>The discriminator in a GAN is simply a classifier. It tries to distinguish real data from the data created by the generator. It could use any network architecture appropriate to the type of data it's classifying. https://developers.google.com/machine-learning/gan/discriminator https://developers.google.com/machine-learning/gan/discrimin...
- robbrown451 3y agoRead the last paragraph. You still have humans, but their input is more akin to a movie reviewer than a movie director/writer/actor/etc. It still takes skill, but it takes a lot less time. RLHF typically employees humans, and that can be time consuming in itself, but less time consuming than creating content. And their efforts can be amplified. If they are actually rating unpaid humans, that is, users, who are willing to provide feedback and are also prompting the system. Plenty of people are happy to do this for free, and some of it happens, just as a byproduct of them doing what they're already doing, creating content and choosing, which comes out good and which one doesn't. Every time I am working through a coding problem with chatGPT, and it makes mistakes and I tell her about those mistakes, it can be learning from that. People can also come up with coding problems that can run and test itself on. As a simple example, I imagine it's trying to write a sorting algorithm. It can also write a testing function simply tests that it is correctly sorted. They can also time its results, they can count how many steps had to do in that sense it can work just like Alpha zero, where there is an objective goal, which is to do it with the least clock cycles, and there's a way to test whether and how well it is a achieving that goal. While that may be a limited number of programming problems that that works for, by practicing on that type of problem it will presumably get better at other types of problems, just like humans do. This is exactly what large language models do, they find a way to objectively test their writing ability, which is by having them predict words and things that they've never seen before. In a sense it's different from actually writing new creative content, but it is practicing skills that you need to tap into when you are creating new content. Interestingly, a lot of people will dismiss them as simply being word predictors, but that's not really what they're doing. They're predicting words when they're training, but when they're actually generating new content, they're not "predicting" words (you can't predict your own decisions, that doesn't make sense), they are choosing words.
- spacecadet 3y agoI got her writing pretty advanced programs that generate fake data sets and self score those data sets. Fun little project to see what would happen.
- riku_iki 3y ago> With Alpha Go, you have a clear objective -- to win a game. How does that work for creative outputs? there are still tons of potentially valuable applications with clear objective: win stock market, create new material or design to maximize some metrics, etc.
- mindwok 3y agoI think similar to humans, creativity will be an emergent behavior as a result of the intelligence needed to pass other tests. Evolution doesn't care about our art, but the capabilities we use to produce it also help us with survival.
- skybrian 3y agoThis isn’t directly about creativity, but I suspect a lot of training will happen in simulated environments. A sandboxed Python interpreter is a good example. There are plenty of programming questions to train on.
- nonameiguess 3y agoThat works fine for purely text-based or digital knowledge domains. So, sure, many types of programming, probably most game play, certainly all video game play, many types of purely creative fictional writing. I don't want to downplay those applications, but the killer breakthrough that the breathless world imagines and has wanted since Turing first talked about this is accurately modeling physical reality. "Invent a better engine" and what not. Without being physically embodied and being able to conduct experiments in the real world, you can't bootstrap that, short of simulating physics from first principles, which is not computationally feasible. You're inherently relying on some quorum of training material produced by embodied sources capable of actually doing science to be factually accurate.
- alexpetralia 3y agoNot dissimilar to how large organizations operate today! Humans operate at the edge collecting sensory data (making measurements, inputting forms, etc.) and the "brain" is a giant management and software apparatus in the middle.
- ooterness 3y agoUnfortunately, I think the current strategies for RLHF are a huge contributor to hallucination / confabulation. In short, they're paying contract workers for quantity, not quality; they don't have time to do independent research or follow up on citations. Unsurprisingly, the LLM optimizes for superficially convincing bullshit.
- quickthrower2 3y ago> Among those who labeled demonstration data for InstructGPT, ~90% have at least a college degree and more than one-third have a master’s degree. Source: https://huyenchip.com/2023/05/02/rlhf.html#demonstration_data https://huyenchip.com/2023/05/02/rlhf.html#demonstration_dat...
- yeck 3y agoControlling for academic experience probably raises the average accuracy of labelling, but by how much? Clearly having a degree will not make you omniscient in your major, let alone other subjects.
- mdekkers 3y agoAre they getting paid on quantity or quality?
- robbrown451 3y agoMost jobs I've known factor in both. I would assume they have processes in place that incentivize quality. Some it is as simple as you have a manager that will fire you if you produce crap. With billions in funding, and bad results causing bad press etc, you think that OpenAI would not have given this a bit of consideration?
- mdekkers 3y ago> I would assume they have processes in place that incentivize quality. > you think that OpenAI would not have given this a bit of consideration? Those are just assumptions though. The issue is not “this was labelled as a shoe, but it’s a car”, the issue is about depth vs superficially, which is harder to verify. See also https://www.theverge.com/features/23764584/ai-artificial-intelligence-data-notation-labor-scale-surge-remotasks-openai-chatbots https://www.theverge.com/features/23764584/ai-artificial-int... for a well-sourced article on the subject.
- dr_dshiv 3y agoWe need alphago for math problems. Anyone know of a project like this?
- rhdunn 3y agoWith AlphaZero there are clear evaliation metrics -- you win, lose, or draw the game given specific rules. With chess, there is even a way of detecting end-game threats via check. The zero human data approach works here because of that, allowing the computer to find optimal strategies. With natural language you don't have that unaided feedback evaluation metric. Especially when given idioms, domain specific terms, etc. This is slow and hardwork because you need to process some text, evaluate and correct that data, retrain and repeat with the next text. You also need to check and correct the existing data, because inconsistencies will compound any errors.