5 ms·
This just confirms I'm not minimally competent in this conversation. Is there a "college freshman" explainer? GPT seems to be doing something incredibly diffe
by Digory 4y ago
This just confirms I'm not minimally competent in this conversation. Is there a "college freshman" explainer?
GPT seems to be doing something incredibly different than prior AI. Is it really a Bayesian "next word" chooser at incredible scale?
- anon291 4y agoYou're looking for the paper 'attention is all you need'. Gpt is not a bayesian next word chooser. It does something different.
- hakuseki 4y agoI think that's not a bad summary, though? Perhaps you would say it is a probabilistic next-token chooser, but that just seems like a very minor distinction.
- anon291 4y agoProbabilistic and bayesian are not identical things. Moreover, GPT the deep-learning model is not a probabilistic next-token chooser. You can envision many different ways to choose the next word based on GPT output. OpenAI's API for GPT is a probabilistic word chooser paired along with GPT. But GPT is the model. It generates a set of probability distributions for the next word, not using a Bayesian process but something entirely different. GPT takes a vector space representation of a sentence and projects it onto some space (we'll call it GPTThink) and then re-projects that space to a new vector space. Then it uses softmax to turn that vector space into a probability distribution. That's not a Bayesian process.
- Digory 4y agoBetter! The last sentence still sounds like "magic," but this is getting closer to my mental comprehension of how you get from BASIC and Python to GPT.
- yunwal 4y agoNot really. Attention is all you need describes a new mechanism used in transformer networks, but the model is still a Bayesian word chooser
- theonemind 4y agoI want to ask GPT to explain how GPT works using a simple metaphor and common language without computer science or AI field terms, where simplicity is more important than correctness. Unfortunately, it would probably make up something extremely plausible sounding and very wrong.
- abi 4y agoI've been enjoying Andrej Karpathy's YouTube series on neural networks: https://www.youtube.com/playlist?list=PLAqhIrjkxbuWI23v9cThsA9GvCAUhRvKZ https://www.youtube.com/playlist?list=PLAqhIrjkxbuWI23v9cThs... It starts from absolute basics and goes slowly. I've only watched about half of it and it has already helped me understand a lot of AI concepts that I see frequently spoken about.