3 ms·
I think the trouble is that "model" is a very general term. If you had a computer doing simulations of artillery shots back in the 50s, then it would have a "mo
by cmdli 3y ago
I think the trouble is that "model" is a very general term. If you had a computer doing simulations of artillery shots back in the 50s, then it would have a "model" of the world in terms of variables tracking projectiles, but this model doesn't generalize to anything else. If a computer does image recognition from the 90s and 2000s to recognize faces, then the computer has a "model" of visual information in the world, but this model only lets it recognize faces.
ChatGPT has a model of all the text information on the internet, but it remains to be seen what the hard limits of this model are. Does this model let it do logic or predict the future well, or will no amount of training give it those abilities? Simply being good in one task doesn't imply a general ability to do everything, or even most of everything. LLM's would simply be the last advancement in a field with a lot of similar advancements.
- famouswaffles 3y ago>ChatGPT has a model of all the text information on the internet, but it remains to be seen what the hard limits of this model are. Before training is complete and loss is maxed, there will be limits on what the "learned so far" model can do that say absolutely nothing about the limits of a perfect(or very close to it) model. It really looks like anything will converge with enough compute. I don't think architecture is particularly important except as "how much compute will this one take?" question. https://nonint.com/2023/06/10/the-it-in-ai-models-is-the-dataset/ https://nonint.com/2023/06/10/the-it-in-ai-models-is-the-dat... >Does this model let it do logic or predict the future well, or will no amount of training give it those abilities? There's nothing special about logic. Basically, any sequence is fair game. It literally does not matter to the machine. Boolformer: Symbolic Regression of Logic Functions with Transformers(https://arxiv.org/abs/2309.12207 https://arxiv.org/abs/2309.12207) That said, GPT-4 can already do logic. It's not perfect but if perfect logic were a requirement then humans cannot do logic either. >Simply being good in one task doesn't imply a general ability to do everything, or even most of everything. It's not one task. It's one modality (text) that a plethora of tasks could be learned in. Coding and playing chess did not suddenly become a single task just because we found the common ground that allows a machine to learn both. The text, image, video and audio data we could feed a transformer will cover anything we care about.