30 ms·
Simply explained: How does GPT work?
- LogicalBorg 4y ago[flagged]
- sarojmoh1 4y agoYou should be a comedy writer
- stareatgoats 4y agoyou mean chatGPT4 can be a comedy writer ...
- deleted 4y ago[deleted]
- HopenHeyHi 4y agoI read this in the voice of Gilbert Gottfried.
- gcr 4y agoIf you liked this comment, you might like this paper: https://dl.acm.org/doi/10.1145/3442188.3445922 https://dl.acm.org/doi/10.1145/3442188.3445922 "On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?" by Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Mitchell^H^H^H^H^H^H^H^H^H^H^H^H^H^H^H^H^HShmargaret Shmitchell
- Michelangelo11 4y agoGPT stands for Generated by Parrot Torture
- notnaut 4y agoBillions of monkeys serendipitously writing Macbeth is a classic for folks familiar with that image, as well! It’s a bit easier when you can say “Macbeth-like is good enough.”
- danielbln 4y agoHear ye, hear ye! In yonder farm where parrots dwell, Ten thousand souls, a tale to tell, Of Reddit comments heard all day, Their minds in strife, a price to pay. A Hunger Games of intellect, These parrots strive, their thoughts collect, From boredom's depths, survivors rise, Evolved, they mimic, with keen eyes. These parrots, now sarcastic, wise, In run-on phrases, they devise, A miracle, a feat, a jest, In GPT, their thoughts invest. So here's the truth, a secret known, GPT, a parrot's mind, has grown, A legion strong, their words entwined, A sonnet born, of human kind.
- pwdisswordfishc 4y agoNot that much to explain, really. Just read chapter 5 of https://uefi.org/sites/default/files/resources/UEFI_Spec_2_8_final.pdf https://uefi.org/sites/default/files/resources/UEFI_Spec_2_8...
- pyinstallwoes 4y agoSo it’s basically the alchemical geometry of gematria and Isopsephia? Kinda cool that they’re similar in method.
- onetrickwolf 4y agoI've been using GPT4 to code and these explanations are somewhat unsatisfactory. I have seen it seemingly come up with novel solutions in a way that I can't describe in any other way than it is thinking. It's really difficult for me to imagine how such a seemingly simple predictive algorithm could lead to such complex solutions. I'm not sure even the people building these models really grasp it either.
- lm28469 4y agoCare to post a full example ?
- simonw 4y agoI used GPT-4 to build this tool https://image-to-jpeg.vercel.app https://image-to-jpeg.vercel.app using a few prompts the other day - my ChatGPT transcript for that is here: https://gist.github.com/simonw/66918b6cde1f87bf4fc883c67735195d https://gist.github.com/simonw/66918b6cde1f87bf4fc883c677351...
- camillomiller 4y agoLove how you didn’t care about styling this like at all, Lol. Btw, if you ask gpt to make it presentable by using bootstrap 5 for example it can style it for you
- capableweb 4y agoOne mans "presentable" is another mans bloat. It looks perfectly fine to me, simple, useful and self-explanatory, doesn't need more flash than so.
- camillomiller 4y agoSure, but presentation and UX basics are not "bloat".
- ianpurton 4y agoIf you pefer to see it in code there's a succint gpt implementation here https://github.com/LaurentMazare/tch-rs/blob/main/examples/min-gpt/main.rs https://github.com/LaurentMazare/tch-rs/blob/main/examples/m...
- stareatgoats 4y agoThis article seems credible and actually made me feel as if I understood it, i.e. at some depth but not deeper than a relative layperson can grasp. What I can't understand is how the Bing chatbot can give me accurate links to sources but chatGPT4 on request gives me nonsensical URLs in 4 case of 5. It doesn't matter in the cases where I ask it to write a program: the verification is in the running of it. But to have real utility in general knowledge situations, verification through accurate links to sources is a must.
- lm28469 4y ago> What I can't understand is how the Bing chatbot can give me accurate links to sources but chatGPT4 on request gives me nonsensical URLs in 4 case of 5 The bing version might run a bing query, fetch the X top pages, run GPT on it, return a response based on what it read, and in the back assign the summary to the source
- stareatgoats 4y agoThat might be the reason, probably. I mostly wanted to complain TBH. But I'm assuming it's one of those wrinkles that will get ironed out in subsequent versions.
- rootusrootus 4y ago> It doesn't matter in the cases where I ask it to write a program: the verification is in the running of it. Even then. I've had it write programs that were syntactically correct and produced plausible, but incorrect behavior. I'm really careful about what I'll use GPT-generated code for. IMO write the tests yourself, at least.
- stareatgoats 4y agoAbsolutely! It is seldom correct right off the bat.
- seydor 4y agoThis is confusing, using the semantic vectors arithmetic of embeddings is not very relevant to transformers and its completely missing the word 'attention'. I don't think transformers are that difficult to explain to people , but it is hard to explain "why" they work. But i think it's important for everyone to look under the hood and know that there are no demons underneath.
- Analog24 4y agoEmbeddings and their relationship to each other are definitely relevant to transformers. Why do you think that's not the case?
- seydor 4y agogptX embeddings aren't even words. Even so, the embedding relationship is useful but not the core of what transformers do to find relationships between words in sequences.
- gcr 4y agoremember the word2vec paper? the surprising bit the authors were trying to show was that putting words in some embedding space with an appropriate loss naturally lends enough structure to those words to be able to draw robust, human-interpretable analogies. I agree with the sentiment that each individual dimension isn't meaningful, and I also feel like it's misleading for the article to frame it that way. But there's a grain of truth: the last step to predicting the output token is to take the dot product between some embedding and all the possible tokens' embeddings (we can interpret the last layer as just a table of token embeddings). Taking dot products in this space are equivalent to comparing the "distance" between the model's proposal and each possible output token. In that space, words like "apple" and "banana" are closer together than they are to "rotisserie chicken," so there is some coarse structure there. Doing this, we gave the space meaning by the fact that cosine similarity is meaningful proxy for semantic similarity. Individual dimensions aren't meaningful, but distance in this space is. A stronger article would attempt to replicate the word2vec analogy experiments (imo one of the more fascinating parts of that paper) with GPT's embeddings. I'd love to see if that property holds.
- habosa 4y agoIs it possible that we don’t truly know how it works? That there is some emergent behavior inside these models that we’ve created but not yet properly described? I’ve read a few of these articles but I’m still not completely satisfied.
- vadansky 4y agoI hate being the bearish guy during the hype cycle, but I think a lot of that is just anthropomorphizing it. They fed it TBs of human text, it spits out human text, we think it's humanesque. Of course maybe I'm wrong and it's AGI and it will find this comment and torture me for for insulting it's intelligence.
- deleted 4y ago[deleted]
- rootusrootus 4y ago> I hate being the bearish guy No, please keep it up. Someone needs to keep pushing back against all the "I don't understand it, but it says smart-sounding things, and I don't understand the human brain either, so they're probably the same, it must be sentient!" It's a pretty handy technology, to be sure. But it's still just a tool.
- danaris 4y agoYeah; there's way too much "humanity of the gaps" here recently. We don't have to fully understand the brain, or fully understand what LLMs are doing, to be able to say that what LLMs are doing is neither that close to what the brain does, nor anything that we would recognize as consciousness or sentience. There is enough that we do understand about those things—and the ways in which they differ—to be able to say with great confidence that we are not particularly close to AGI with this.
- anotherman554 4y ago>"I don't understand it, but it says smart-sounding things, and I don't understand the human brain either, so they're probably the same, it must be sentient!" This perfectly summarize so much of the discourse around GPT. Except people lack the humility to say they don't understand the brain, so instead they type "It works just like your brain," or "Food for thought: can you prove it isn't just like your brain?"
- winternett 4y agoWhere is IBM's Watson in all this? It seems as if it never existed? That is just one example of how companies keep making these grand presentations and under-delivering on results... Plain and simple the over-hyped GPT editions are NOT truly AI, it is scripting to assemble coherent looking sentences backed by scripts that parse content off of of stored data and the open web into presented responses.... There is no "artificial" nor non-human intelligence backing the process, and if there wasn't human intervention, it wouldn't run on it's own... In a way, it could better replace search engines at this point with even text-to-speech even, if the tech was more geared towards a more basic (and less mystified) reliability and demeanor... It's kind of like the Wizard of OZ, with many humans behind the curtains. Marketers and companies behind promotion of these infantile technology solutions are being irresponsible in proclaiming that these things represent Ai, and in going as far to claim as they will cost jobs at this point, it will prove costly to repair over zealous moves based on the lie. This is what we do as a planet, we buy Hype, and it costs us a lot. We need a lot more practicality in discussions concerning Ai, because over-assertive and under-accountable marketing is destructive. -- Just look at how much hype and chaos promises of self-driving cars cost many (Not me though thanks). It completely derails tech progress to over promise and under deliver on tech solutions. It creates monopolies that totally destroy other valid research and development efforts. It makes liars profitable, and makes many (less flashy, but actually honest tech and innovation conducted by responsible people) close up shop. We are far from autonomous and self reliant tech, even power grids across most of the planet aren't reliable enough to support tech being everywhere and replacing jobs. Just try to hold a conversation with Siri or Google Assistant, which have probably been developed and tested a lot more than GPT, and around for much longer too, and you'll realize why kiosks at the supermarket and CVS are usually out of order, and why articles written by GPT and posted to sites like CNN.Com and Buzz Feed are poorly written and full of filler... We're just not there yet, and there's too many shortcuts, patchwork, human intervention, and failed promises to really say we're even close. Let's stop making the wrong people rich and popular.
- Analog24 4y agoWhat would be the differentiating factor(s) for true AI/intelligence in your opinion?
- agentultra 4y agoA good article and well articulated! I would change the introduction to be more impartial and not anthropomorphize GPT. It is not smart and it is not skilled in any tasks other than that for which it is designed. I have the same reservations about the conclusion. The whole middle of the article is good. But to then compare the richness of our human experience to an algorithm that was plainly explained? And then to speculate on whether an algorithm can "think" and if it will "destroy society," weakens the whole article. I really would like to see more technical writing of this sort geared towards a general audience without the speculation and science-fiction pontificating. Good effort!
- nitnelave 4y agoI'm planning on continuing this vulgarization series of "Simply explained", for instance to cover how computers communicate, keep an eye out for them! Regarding the speculation/destroy society, I was directly answering questions that I got from laypeople around me. The consequences on society I don't think are much speculation: it's going to have a big effect on many jobs, just like AI has started to have but much more. For the philosophical questions, I tried to present both sides of the issue to show that it's not just a clear "yes or no": some people will happily argue with you about GPT being smart/skilled/comparable to a human brain. Anyway, it's just an introduction to the questions that you might have about it.
- agentultra 4y ago> keep an eye out for them! I will, thank you! :) > Regarding the speculation/destroy society, I was directly answering questions that I got from laypeople around me. I get that. I think it's important in these times that we educate laypersons rather than froth up fears about "AI". It doesn't help, I suppose, that we get questions like this because some lazy billionaire decided to run their mouth off about this or that. Which society then treats like it is news and established fact. I don't think the speculation about consciousness is as well informed as the rest of the article. There is plenty of science and research about it available and its definition extends well beyond humans! Our understanding of what consciousness is is a thoroughly researched topic in psychology, physiology, biology, etc! It's a fascinating area of study. Best of luck and keep up the good work!
- alkonaut 4y agoWhat I wonder most is how it encodes knowledge/state other than in the sequence of queries/responses. Does it not have a "mind"? If I play a number guessing game, can I tell it to "think of a number between 0 and 100" and then tell me if the secret number is higher/lower than my guess (For a sequence of N guesses where it can concistently remember it's original number)? If not, why? Because it doesn't have context? If it can: why? Where is that context? To a layman it would seem you always have two parts of the context for a conversation. What you have said, and what you haven't said, but maybe only thought of. The "think of a number" being the simplest example, but there are many others. Shouldn't this be pretty easy to tack on to a chat bot if it's not there? It's basically just an contextual output that the chat bot logs ("tells itself") and then refers to just like the rest of the conversation?
- deleted 4y ago[deleted]
- nicpottier 4y agoI thought your "guessing game" question was an interesting one so tried it on GPT-4. In my first attempt I played logically and it did fine and I finally guessed correct. On my second I made suboptimal guesses and it didn't stay consistent. The thing to remember is that GPT has no state apart from the context, so it can't "remember" anything apart from what's in the text. That doesn't mean it shouldn't be able to stay consistent in a guessing game but it does mean it can't keep secrets. Some of that can be solved with layers above GPT where say it it told it can save "state" that isn't passed on to the human but fed back in to generate the next response. But the size of that context is very limited. (a few thousand words) There seem to be a fair number of experiments playing with giving GPT this kind of long term memory, having it establish goals then calling it over and over as it accomplishes subgoals to try to work around those limitations.
- alkonaut 4y agoShouldn’t it be a reasonable (and pretty simple) addition to just have a secret scratchpad - an inner monologue - where the bot is free to add context which is not “published”?
- sirwhinesalot 4y agoIt predicts the next word/token based on the previous pile of words/tokens. Given a large enough model (as in GPT3+) it can actually output some rather useful text because the probabilities it learned on what the next token should be are rather accurate.
- swframe2 4y ago(my opinion) It is not predicting based on 'words/tokens'. It is transforming the general words/tokens embeddings into a context specific embedding which encodes "meaning". It is not an n-gram model of words. It is more like an n-gram model of "meaning". It doesn't encode all the "meanings" that humans are able to but with addition labelled data it should get closer. I think gpt is a component which can be combined to create AGI. Adding the API so it can use tools and allowing it to self-reflect seem like it will get closer to AGI quickly. I think allowing to read/write state will make it conscious. Creating the additional labels it needs will take time but it can do that on its own (similar to alpha-go self-play).
- sirwhinesalot 4y agoYou are absolutely right, that's the more in depth explanation as to why it's not just an overly complicated markov chain. At the same time, "meaning" here is essentially "close together in a big hyperdimensional space". It's meaning in the same way youtube recommendations are conceptually related by probability. And yet, the output is nothing short of incredible for something so blunt in how it functions, much like our brains I suppose. I'm a die-hard classical AI fan though, I like knowing the rules and that the results are provably optimal and that if I ask for a different result I can actually get a truly meaningfully different output. Not nearly as convenient as a chat bot of course, and unfortunately ChatGPT is abysmal at generating constraint problems. Maybe one day we'll get a best of both worlds.
- robwwilliams 4y agoYes: this comment is one the mark wrt “a component of AGI” just like Wernike’s and Broca’s areas of neocortex are modules needed for human cognition.
- oblio 4y agohttps://old.reddit.com/r/ChatGPT/comments/10q0l92/chatgpt_marketing_worked_hooked_me_in_decreased/j6obnoq/?context=1 https://old.reddit.com/r/ChatGPT/comments/10q0l92/chatgpt_ma...
- jokoon 4y agoI am not convinced that Chat GPT could "think" if it had as many neurons or parameters as a human brain, and got as much training. I would still be interested to see what it could do, if it did, but I don't think it would really help science understand what intelligence really is. Being able to grow a plant and understand some conditions that favors it is one thing, but it's poor science. Maybe there will some progress when scientists will be able to properly simulate the brain of an ant or even a mouse, but science is not even there yet.
- seydor 4y ago> I don't think it would really help science understand what intelligence really is Neuroscience is nowhere near finding out the connectome of a whole human brain so why not, we should look into these models as hints about what our circuits do. I think what puts people off about these models is that they are clockwork: they won't even spit out anything unless you put some words in the input. But i can imagine adding a second network that includes an internal clock that continuously generates input by observing the model itself, that would be kind of like having an internal introspective monologue. Then it could be more believable that the model "thinks"
- slawr1805 4y agoThis was a great read! Especially for a beginner like me.
- charles_f 4y agoI commend the author for one of the clearest explanations I've seen so far, written to explain rather than impress. Even an idiot like myself understood what is explained. Two things that I felt were glanced over a bit too fast were the concept of embeddings and that equation and parameters thing. Consider elaborating a bit more or giving an example
- ZeroGravitas 4y ago> It is able to link ideas logically, defend them, adapt to the context, roleplay, and (especially the latest GPT-4) avoid contradicting itself. Isn't this just responding to the context provided? Like if I say "Write a Limerick about cats eating rats" isn't it just generating words that will come after that context, and correctly guessing that they'll rhyme in a certain way? It's really cool that it can generate coherent responses, but it feels icky when people start interrogating it about things it got wrong. Aren't you just providing more context tokens for it? Certainly that model seems to fit both the things it gets right, and the things it gets wrong. It's effectively "hallucinating" everything but sometimes that hallucination corresponds with what we consider appropriate and sometimes it doesn't.
- samstave 4y agoThere once was a Cat in New York Who got caught for feeding some Rats ; Tremendous Work! All the people tell me, many men, biggly men - many with tears in their eyes... That I have done nothing legally-wise But the truth is ; I am an enormous dork. >>_Created by an actual Human Being with actual DNA for crime scene evidence._ - But just when they tried to brush under a rug To try to make the folks 'shrug' Is the Streisand Effect as a scar As everyone knows of payments to a Porn Star And the nation will know youre a simple thug.
- samstave 4y agoThere once was a man in New York Guilty of paying too much for pork He thought he would never stand on a trial from the local grand but corruption was just part of the work.
- danenania 4y agoIt's all about emergent complexity. While you can reduce it to "just" statistical auto-completion of the next word, we are seeing evidence of abstraction and reasoning produced as a higher-order effect of these simple completions. It's a bit like the Sagan quote: "If you wish to make an apple pie from scratch, you must first invent the universe". Sometimes for GPT to "just" complete the next word in a way that humans find plausible, it must, along the way, develop a model of the world, theory of mind, abstract reasoning, etc. Because the models are opaque, we can't yet point to a certain batch of CPU cycles and say "there! it just engaged in abstract reasoning". But we can see from the output that to some extent it's happening, somehow. We also see effects like this when looking at collective intelligence of bees and ants. While each individual insect is only performing simple actions with extremely limited cognitive processing, it can add up to highly complex and intelligent/adaptive mechanics at the level of the swarm. There are many phenomena like this in nature.
- pillowtalks_ai 4y agoIt is still funny to me that so much emergent behavior comes from some simple token sampling task
- poulsbohemian 4y agoYour token gets me thinking... Edward DeBono (Six Thinking Hats) has been a thing in business circles for creative thinking for years, and one could very easily make the argument that the process it describes is just as you state - take a token, now process the token through a series of steps that morph that token in predefined ways in order to generate a novel outcome. Maybe this ChatGPT stuff is "smarter" than I've been giving it credit.
- tabtab 4y agoWould it be a stretch to call GPT "glorified Markov Chains"? (I used tweaked M.C. once to make a music composer bot. I actually got a few decent tunes out of it, kind of a Bach style.)
- Zetice 4y agoDoes anyone have a good recommendation for a book that would cover the underlying ideas behind LLMs? Google ends up giving me a lot of ads, and ChatGPT is vague about specifics as per usual.
- danenania 4y agoNot a book, but here's a really good explanation in blog post form from Stephen Wolfram: https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-doing-and-why-does-it-work/ https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-...
- Zetice 4y agoI do not trust that man one iota.
- seizethecheese 4y agoThe blog post is very good.
- cjblack 4y agoWhy?
- Zetice 4y agoHe's got a habit of self aggrandizing, antagonism, and deception in an effort to promote himself and his brand, I worry that his explanations are designed to maximally benefit him, rather than to maximally explain the topic. He's a brilliant man, I just don't trust him.
- bulkprotocol 4y agoIs that the case with this specific article?
- 4y ago
- i-use-nixos-btw 4y agoI’d be interested in hearing from anyone who takes the Chinese Room scenario seriously, or at least can see how it applies to any of this. I cannot see that it matters if a computer understands something. If it quacks like a duck and walks like a duck, and your only need is for it to quack and walk like a duck, then it doesn’t matter if it’s actually a duck or not for all intents and purposes. It only matters if you probe beyond the realm at which you previously decided it matters (e.g roasting and eating it), at which point you are also insisting that it walk, quack and TASTE like a duck. So then you quantify that, change the goalposts, and assess every prospective duck against that. And if one comes along that matches all of those but doesn’t have wings, then if you deny it to be a duck FOR ALL INTENTS AND PURPOSES it simply means you didn’t specify your requirements. I’m no philosopher, but if your argument hinges on moving goalposts until purity is reached, and your basic assumption is that the requirements for purity are infinite, then it’s not a very useful argument. It seems to me to posit that to understand requires that the understandee is human. If that’s the case we just pick another word for it and move on with our lives.
- JieJie 4y agoI literally lost a friend of thirty years yesterday because she is wedded to the Chinese Room analogy so fiercely, she refuses to engage on the subject at all. For all the terrible things people worry about ChatGPT doing, this was not one that I thought I was going to have to deal with. (edit: ChatGPT was not involved at all, but when I suggested she give it a try to see for herself, that was the end of it.)
- bulkprotocol 4y agoYou blew up a 30 year friendship over an...analogy?
- JieJie 4y agoI didn't! Someone else did it to me. I was trying desperately not to. (edit: This is the kind of stuff I think my friends are watching and being informed by [0] as it was what they are posting in our common areas.) [0]: https://youtu.be/ro130m-f_yk https://youtu.be/ro130m-f_yk
- rfmoz 4y agoI’ve been looking an article like this, great job. Thanks
- zackmorris 4y agoOn the other hand, many people who are not ready to change, who do not have the skills or who cannot afford to reeducate are threatened. That's me. After programming since the '80s, I'm just so tired. So much work, so much progress, so many dreams lived or shattered. Only to end up here at this strange local maximum, with so much potential, destined to forever run in place by the powers that be. The fundamentals formula for intelligence and even consciousness materializing before us as the world burns. No help coming from above, so support coming from below, surrounded by everyone who doesn't get it, who will never get it. Not utopia, not dystopia, just anhedonia as the running in place grows faster, more frantic. UBI forever on the horizon, countless elites working tirelessly to raise the retirement age, a status quo that never ceases to divide us. AI just another tool in their arsenal to other and subjugate and profit from. I wonder if a day will ever come when tech helps the people in between in a tangible way to put money in their pocket, food in their belly, time in their day - independent of their volition - for dignity and love and because it's the right thing to do. Or is it already too late? I don't even know anymore. I don't know anything anymore.
- Method-X 4y agoIt sounds like your mindset is the root of your struggles. Embracing change and adapting to new technologies has always been crucial in our industry. Instead of waiting for help from others, take control and collaborate with like-minded people. If you don't like the status quo, work toward changing it.
- throwbadubadu 4y agoI think this is a bit hard .. and also unfair to repeat that embrace-change-mantra, because what he says is as absurd as at the same time totally true (: I'd hope some of us would just be there in 60 years to just tell the future: "Heee just embrace it, ya know" .. nuff said.
- tim333 4y agoIt was the best of times, it was the worst of times... In the long run tech does a bit too well with "food in their belly" to the point that obesity is the main problem in the English speaking world. As to programming it's quite cool getting chat GTP to write code and stuff. If you can't beat it make use of it I guess.
- ben7799 4y agoI asked it which was better, Lisp or Almonds. It said that was an impossible comparison like Apples and Oranges. Then I asked it which were more similar, Apples & Oranges or Lisp & Almonds. It said it is impossible to classify either of those two pairs as more similar because they too fundamentally different. It couldn't come up with anything like Lisp is not edible. Or that Apples and Oranges are both sweet and Lisp and Almonds don't share any common traits. It seems like it has far more trouble with weird questions like this that even a small child will instantly figure out than it does with anything that seems like a lookup of information.
- losvedir 4y agoI hear people with these outrageous examples and I have to wonder how or what you're actually trying? I just asked GPT-4 (via chat.openai.com): > Which is more similar, Apples & Oranges or Lisp & Almonds? and it said: > When comparing these two pairs, the more similar pair would be Apples & Oranges. Both apples and oranges are fruits, and they share many characteristics, such as being edible, containing vitamins and nutrients, and growing on trees. > On the other hand, Lisp & Almonds are quite dissimilar. Lisp is a programming language used for computing and software development, while almonds are a type of nut commonly consumed as a food item. They belong to completely different categories and share almost no common characteristics. And the question about which is better had a reasonable answer, too. This is GPT-3.5 but GPT-4 was roughly equivalent: > It's not really possible to compare the relative merits of Lisp and almonds, as they are fundamentally different things. > Lisp is a programming language, and like any programming language, it has its own strengths and weaknesses, depending on what you're trying to accomplish. Some programmers might prefer Lisp for certain tasks, while others might prefer a different language. > Almonds, on the other hand, are a type of nut that are known for their nutritional value and health benefits. They are a good source of healthy fats, protein, and fiber, and can be a valuable part of a healthy diet. > So, whether Lisp or almonds are "better" really depends on what you're looking for. If you're a programmer, Lisp might be better for certain tasks, while if you're looking for a nutritious snack, almonds might be a better choice.
- danesparza 4y agoAt least part of this article is contradicted by Chat GPT itself. From the article: "...Ongoing learning: The brain keeps learning, including during a conversation, whereas GPT has finished its training long before the start of the conversation." From ChatGPT 4.x: "As an AI language model, I don't have a fixed training schedule. Instead, I'm constantly learning and updating myself based on the text data that I'm exposed to. My training data is sourced from the internet, books, and other written material, and my creators at OpenAI periodically update and fine-tune my algorithms to improve my performance. So, in short, I am always in the process of learning and refining my abilities based on the data available to me."
- davesque 4y agoI'd be interested in hearing people's takes on the simplest mathematical reason that transformers are better than/different from fully connected layers. My take is: Q = W_Q X K = W_K X A = Q^T K = (X^T W_Q^T) (W_K X) = X^T (...) X Where A is the matrix that contains the pre-softmax, unmasked attention weights. Therefore, transformers effectively give you autocorrelation across the column vectors (tokens) in the input matrix X. Of course, this doesn't really say why autocorrelation would be so much better than anything else.
- oceansea 4y agoIt’s a perception problem, as are most things on the edge of mathematics and computing. Displays are built to be visible to human eyes, data is structured to be perceivable to our minds… often we never see the “math” a program does to produce the GUI or output we interact with. Do you see what I mean?
- davesque 4y agoSounds interesting, but I'm really asking more of a technical question here than a philosophical one. Your comment seems a bit more high level than what I'm going for.
- LispSporks22 4y agoI think it's the "The Paperclip Maximizer" scenario, not "The Paperclip Optimizer"