4 ms·
One thing OpenAI has now that it didn't have 4 years ago is a lot more compute power at its disposal. Sam Altman has already said "I think we're at the end of t
by gexla 3y ago
One thing OpenAI has now that it didn't have 4 years ago is a lot more compute power at its disposal. Sam Altman has already said "I think we're at the end of the era where it's going to be these, like, giant, giant models." If that's actually true, then the GPTX tech has largely hit a wall in which throwing more compute at it won't get the same increase in capability. Bill Gates predicted that GTP5 won't be much better than GPT4. So, any "breakthrough" really could be incremental rather than game changing.
- visarga 3y agoI think there are two main source of learning for AI - the web scrape datasets, which contain our historical experiences and communications, and AI feedback generated from deployed agents. The web text is almost exhausted, or we can't scale it 100x more, but the feedback is just starting to ramp up. Every day millions of chat sessions are recorded, and they are exceptional training examples. They would contain the kind of errors LLMs do, and the kind of demands people have, and include a human reaction to each LLM message. The OpenAI move to create "GPTs" is showing they are actively working on improving the feedback signals by empowering the LLM with RAG, code execution and API access. In such a setup it is possible to use a model at level N to generate data at level N+1. The keyword here is learning from feedback, which aligns with recent talk of using RL methods like AlphaZero with LLMs. A RL agent would create its own data as it goes. I think progress will be gradual, as we need to wait for the world to produce the learning feedback signal. Of course in domains where we can speed that up, AI will progress faster. Interesting thought: by making LLMs available to the public, they are going to assist people in many ways and create effects that will percolate in the next training set: LLM inference -> text -> effects in the world -> text -> LLM training. So there is already an implicit feedback loop when we retrain the base models. GPT-5 will train on data from a world influenced and shaped with GPT-4.
- oceanplexian 3y agoThe corollary of this is optimistic though. If increasingly large models don't make a big difference, this is hinting at the fact that data quality is a lot more important than the raw quantity of data. This is good news for open source models, because it would be possible to run or train viable models on less expensive hardware, and having billions of dollars in GPUs isn't as much of a moat as they are suggesting.