15 ms·
Ask HN: How does ChatGPT work?
I'd love a recap of the tech for someone that remembers how ANNs work but not transformers (ELI5?). Why is ChatGPT so much better, too? and how big of a weight network are we talking about that it retains such a diverse knowledge on things?
- catsforai 4y ago
- dtagames 4y agoChatGPT is really very simple. Imagine you could analyze a million books and identify all the words within them -- not the meanings of the words, just the actual letters they contain. Now, when someone asks you about the history of France (or why the sky is blue), you could simply pluck out of your library the most common strings of word that seem to follow the words that were in your question! It's like a kid in the 80's who thinks the answer to an essay question is to copy it from an encyclopedia, only the "encyclopedia" is very large and contains multiple sources. So, the big take away needs to be that there is absolutely no understanding, no cognizance of any kind, no language comprehension going on. The answers look good because they contain all the same words as the most popular answers people have already written which the system scanned. So ChatGPT turns out to great for parsing and summarizing documents, if that's something you need. But, since it doesn't know fact from fiction, it cannot apply logic or math, and it cannot perform reasoning or analysis, it's not good for finding out facts or discerning truth. Another great failing of LLM software is that the user being spoken to is generic. The answers are not modeled for you, they're the same models for everyone. But a human teacher does their job by being exactly the opposite of this -- someone who is finely tuned to the needs and understandings of their audience. A good journalist or writer does the same.
- phillipcarter 4y agoI don't think this is an accurate depiction and I think there's a lot more going on there than you're saying there is. The models also involve Codex, which is able to effectively do pattern matching with more context you give it so that it can emit novel code. I've seen this happen numerous times with Copilot, where it works very well the more open documents you have, especially if those documents are related. True, it doesn't have a full semantic understanding of your codebase, but it's good enough to usually generate the right code once it's been "seeded" with something like an updated type definition that you then need to plumb through various parts of a codebase.
- dtagames 4y agoYes, I use Copilot, too for that very reason. But we need to be very careful about words like "semantic" and "understanding" as the method is neither. We'll get the best use out of these technologies if we don't ascribe to them magical qualities they don't have. Poke around. You'll find out it just statistical math with tokens (letters and punctuation). No meaning is ascribed to anything.
- phillipcarter 4y agoI...did say that it doesn't have a full semantic understanding of your code. Copilot uses tree-sitter under the covers to be able to figure out several things that you can get by analyzing syntax alone, which is actually quite far. It lets you identify what kinds of declarations (e.g., a type) and what kind of expressions (e.g., pattern matching on that type) exist, provided there's a grammar for it, and then lets that influence the suggested code. This isn't perfect because indeed, you need a full language service to actually understand things, especially in resolving ambiguities in symbols (like shadowing in some languages), but, like I said...it's good enough.
- dtagames 4y agoYes, Copilot has the advantage that programming language syntax is highly constrained and regularized, allowing it to slot-in your own variable names and functions into new code as you write it. This is the best feature of Copilot. It's not part of ChatGPT because that's not meant for writing code. And I wouldn't call it "understanding." But, it is useful.
- krackers 4y ago> you could simply pluck out of your library the most common strings of word that seem to follow the words that were in your question! This is not sufficient to explain how LLMs are able to synthesize novel, coherent poems or song lyrics. What you're describing seems closer to a markov model. So far I have yet to see a good explanatin of _why_ transformer models seem to have this emergent behavior as you scale it up. It's easy to see by construction that the stacked attention layers will work well to predict the masked token in "Capital of France is [MASK]", or how you can throw an additional head on top for sentiment analysis (which I recall was implemented by literally adding a dummy token at the start into which information about the entire sentence gets embedded). But it's not obvious (to me at least) that this would somehow generalize into being able to do the things chatGPT does.
- dtagames 4y agoAh, but in fact it does. Code and poetry seem different to people but tokens are tokens to the computer, and it knows not one from the other. The reason you get a poetry-style answer when you ask for one is that the mix of words and styles can be taken from different parts of the corpus. This is how you can have it write (bad) poetry about something for which only prose was scanned.
- nullish_signal 4y agoI think that the levels of Associativity between Linguistic Terms goes beyond a mere "mix of words and styles". When pulling on one part of the Prompt causes portions of the Output to shift or vanish, or change genre or word choice or composition or focal point or focal distance - - - then, that's the very opposite of "mix" -That's "Stable Diffusion" in text2img, or the "Reversal of Entropy" in Physics terms.. I'm still blown away that computers have become inage "Imaginers", decades after Humans first started doing 3D through Computers. It's really all quite unbelievable.
- maxbond 4y agoI'd propose that you come up with a prompt to test your hypothesis and try it out. I think seeing is believing here, I don't know how ChatGPT works but after experimenting with it I am sure it's not simply looking for runs of similar words to the prompt in it's training set. Here's an example to consider: https://news.ycombinator.com/item?id=33874110 https://news.ycombinator.com/item?id=33874110
- nullish_signal 4y agoOP question mentions specific network types, while reply uses a vague "million-book library" metaphor, but at least mentions LLM? Also arbitrarily sorts human { teacher, journalist, writer } as "people who do custom work with each client" which makes a bit of sense for a Teacher in a small-class environment. But Writers and Journalists write at SCALE, and while each may speak with gusto to individuals just as well as they do their job, tailoring custom content to individuals is not their job. Not to mention that individual-tailored model files are very handy with AI, share-able as those diles may be. I call ChatGPT on you!!! And also do not know anything about your question, OP...
- dtagames 4y agoI wouldn't deign to have ChatGPT write my comments. They're all me! Obviously, an author's audience can include more than one person. But every good piece of human writing is tailored for some audience. That was my point. You've heard the expression, "Know your audience?" ChatGPT cannot.
- nullish_signal 4y agoChatGPT's "audience" can range from "hate speech for anti-bias training" to "Teenage Mutant Ninja Turtles Deliver Impassioned Speech to The World" I would also like to use the No True Scotsman fallacy, and assert that no Human can ever really "know" or "write for" anybody's perfect preferences.
- dtagames 4y agoI wasn't asking for perfect preferences. I was asking for minimal cognizance, which the tool cannot provide. ChatGPT is more like a parrot than tech dreamers like to admit.
- nullish_signal 4y agoI'll grant you that ChatGPT is more Parrot than All-Seeing Oracle, but no Parrot (or even Google) will take nginx errors for input, and give me back chmod and chown commands to fix the permission issues. AND! the commands it gave me were tailored to match my user and any directories/websites i told it I was working on ! !
- the_gipsy 4y ago> So, the big take away needs to be that there is absolutely no understanding, no cognizance of any kind, no language comprehension going on. Okay, but what is understanding, cognizance, comprehension? And so on for further precision of definitions, until you reach that vague nebulous area that is consciousness. We simply haven't figured that out, period. So we cannot really say that this AI does not have those qualities, unless its output is obviously showing it. Which ChatGPT does not, most of the time.
- deleted 4y ago[deleted]
- 6nf 4y ago> it cannot apply logic Have you tried?
- pushedx 4y agoAsk ChatGPT to write a song lyric for your favorite band in the style of your favorite poet, with a specific unusual topic. Then claim that the model is “really very simple”
- dtagames 4y agoAt the end of the day, all computing is very simple. What's impressing you here is the size of the dataset and the speed of the retrieval, brought about by advances in hardware. Also, we like the well-formed writing which is created by filtering and massaging the parroted text into a new set of sentences. It is impressive, and people who don't know how it really works will think it's capable of all kinds of things it's not. Of course, this has been true forever. Just ask anyone who writes code about what non-devs think software can do.
- pushedx 4y agoHave you used ChatGPT? Your understanding of GPTs and neural nets in general is consistently flawed. You are describing a Markov model or at best an SVM model. Would you make the same types of claims about Stable Diffusion, that it is “piecing together pieces of existing images?” That is not how these models work. Similar to the temporal memory that ChatGPT will appear to have during a conversation, “inpainting” with Stable Diffusion produces entirely novel output with context. Have you tried it out?
- dtagames 4y agoI have and I wrote extensively about it after seeing the unpopularity of my original comment. There's a link to my long Medium story about it in another comment. Would I say that DALL-E, etc. also piece together existing images? Yes, of course, since that is the only way that technically it could work. All of computing is "input-process-output." There is no "create" step in any of computing. The pieces that DALL-E pieces together are pixels of a certain color and brightness and placement relative to other pixels. Together, we perceive this as an image. If they weren't arranged that way, we'd just see noise or gray. The same is true with ChatGPT. Unless they're put in a particular order, words just make word salad. It is the order that gives them their perceived meaning in a sentence. ChatGPT crafts "new" sentences by rearranging words from similar sentences it has already seen, just like DALL-E creates new paintings by rearranging pixels from paintings it has already seen. A recent video from Noam Chomsky explains this from his perspectives on learning and language: https://www.youtube.com/watch?v=PBdZi_JtV4c https://www.youtube.com/watch?v=PBdZi_JtV4c
- stevenhuang 4y ago> Another great failing of LLM software is that the user being spoken to is generic. The answers are not modeled for you, they're the same models for everyone. But a human teacher does their job by being exactly the opposite of this -- someone who is finely tuned to the needs and understandings of their audience. A good journalist or writer does the same. ChatGPT can do this if you just ask. E.g., explain as if I'm 5 and it will use simple words and phrases. Explain like I'm an expert and it will use technical jargon and expect a deeper shared knowledge with which it can draw from. You shouldn't write so authoritatively when clearly you haven't explored its capabilities.
- rakejake 4y agoThis is true but does not refute OP's point. There is an entire subreddit for ELI5 that is probably part of its training data. So if you ask for ELI5, that pattern IS part of the training corpus.
- weird-eye-issue 4y agoIt is more complex than that, it isn't just scraping /r/ELI5 lol Ask it a question about JavaScript and tell it to respond as a 1920s gangster and it will happily oblige and do a great job. And just to make sure we are on the same page, JavaScript was not available in the 20s
- dtagames 4y agoIt is just scraping and reorganizing words while including words from your prompt in the equation. I already mentioned that the style (gangster) and vocabulary (JavaScript) can come from different scanned documents. Cool? Yes. Is it learning, or understanding? No.
- weird-eye-issue 4y agoIf it's not learning then how can you tell it abstractly why it's code is wrong then have it fix it? Without telling it specifically "change this to this"
- dtagames 4y agoI wrote this[0] to better explain my wildly unpopular answer above. Everything is absolutely correct here, but it's not a popular view right now. Often very smart people like those on HN can be fooled by the "magic" of tech that might not really exist in the way they think it does. [0] https://medium.com/gitconnected/behind-the-curtain-understanding-the-magic-of-chatgpt-3bbd23f0fbb3 https://medium.com/gitconnected/behind-the-curtain-understan...
- xenospn 4y agoChatGPT is a variant of the popular GPT-3 language model, specifically designed for chatbot applications. It uses a combination of deep learning and natural language processing techniques to generate human-like responses to text input in a conversation. The way it works is by first pre-training the model on a large corpus of text data, which could include things like social media conversations, movie scripts, books, etc. This allows the model to learn the general structure and patterns of language. Then, when given an input in the form of a question or statement, the model uses its pre-trained knowledge to generate a response. It does this by predicting the next word in the sentence, and then continuing to predict subsequent words until it reaches the end of the response. Overall, the goal of ChatGPT is to enable chatbots to have more natural, human-like conversations with users. (I asked ChatGPT to tell me how it works)
- ramraj07 4y agoMarginally funny, utterly useless and counterproductive is what it is (this particular use of chatGPT that is).
- benatkin 4y agoHere's an answer I wrote myself: It gives people a heavily filtered and throttled interface to access to their language model. --- Edit: I asked chat.openai.com about it: Would "It gives people a heavily filtered and throttled interface to access to their language model." be a fair way to describe chat.openai.com? Yes, that is an accurate description of chat.openai.com. The site provides users with a limited and controlled interface to access a language model, which is a type of artificial intelligence that is capable of processing and generating natural language. This allows users to have conversations with the language model and get responses based on the information it has been trained on.
- seydor 4y agothats only half the truth, in which the AI sneakily omits the role of human reviewers that worked tirelessly to align its output and make it useful. Shame on the AI
- UncleMeat 4y ago
- adjusted 4y agoIt's still transformer underneath, but openai researchers have figured out how to improve it through engineering efforts and improved training data. I believe it's not easy for outsiders without large model pretraning experience like most of us to understand the tunning details.
- bryan0 4y agoI found this description of the GPT-3 transformer architecture useful: https://dugas.ch/artificial_curiosity/GPT_architecture.html https://dugas.ch/artificial_curiosity/GPT_architecture.html Not eli5 but close enough.
- petesergeant 4y agoSo it really is just a fancy Markov chain?
- dotancohen 4y agoNo.
- tarvaina 4y agoA Markov model is a stochastic process with a discrete visible state where the next state depends only on the current state. GPT’s state is neither discrete nor visible so it is not a Markov model.
- miltondts 4y agoWhat I don't understand is where is the memory? How does GPT-3 or ChatGPT remember so much information with just that architecture? It would seem that the maximum it could remember is 2048 words. EDIT: Maybe it's 2048 x 96? Still seems low for what it can do.
- subroutine 4y agoHow does it know when to stop when asked for a description or summary? Sometimes it outputs a few sentences, sometimes a few paragraphs. Does it know how much output it has already provided when deciding on the next token? How does it decide to start a new sentence or paragraph, or if it's 'satisfied' with its current response?
- amilios 4y agoIn the same way that it predicts a given token like 'we' or 'write', it predicts a special end-output token, something analogous to 'EOF'. So at some point as it's generating new tokens and coming up with probability distributions for the next token (conditioned on its own output so far), selecting from this distribution (either by greedily picking the highest probability token or sampling), it will select this 'EOF'-style token, and end its own output. Does that make sense? So to answer your question about knowing how much text it has output so far, not exactly. It is not 'stateful' exactly, it's simply conditioned on previous text and gives a new probability distribution over its vocabulary for the next token. Whether the 'previous text' is its own output, or provided by a human user (a 'prompt'), it isn't necessarily aware of this by default (unless the user vs. system responses are delineated somehow with special tokens for example, which with ChatGPT they may be/are likely to be. My point is that these models don't come with this awareness built in, you have to add it. Fundamentally they simply condition their probability distribution on a piece of existing text).
- akelly 4y agoThe way they went from GPT-3 to ChatGPT is really quite genius. My understanding is that it's something like this: 1. Start with GPT-3, which predicts the next word in some text and is trained on all the text on the internet 2. Take thousands of prompts, generate several responses for each of them, and have human reviewers rank the responses for each prompt from best to worst 3. The GPT model needs a massive amount of training data, it would be cost prohibitive to get enough human feedback to fine tune GPT manually. So you train another model, called the reward model, to predict how the humans will rate each response. Then you train the GPT model against the reward model millions of times 5. Feed a small percentage of the output from that training process back to the human reviewers to continue training the reward model, based on heuristics like reward model uncertainty which predict how helpful the human feedback will be towards improving the reward model 6. Release ChatGPT to the public, and use user feedback like response upvotes/downvotes to further optimize the reward model, while continuing to train ChatGPT against the reward model https://openai.com/blog/chatgpt/ https://openai.com/blog/chatgpt/ https://openai.com/blog/deep-reinforcement-learning-from-human-preferences/ https://openai.com/blog/deep-reinforcement-learning-from-hum...
- amelius 4y agoThat's not genius, that's called unsupervised learning and it is an entire subfield.
- dsr3 4y agoI think number 3 is a description of Generative Adversarial Network (GAN).
- amelius 4y agoOk, regardless, it is not really new. ML researchers are doing these kinds of things all the time. By the way, according to some people, GANs are also a kind of unsupervised learning: https://stackoverflow.com/questions/44445778/are-gans-unsupervised-or-supervised https://stackoverflow.com/questions/44445778/are-gans-unsupe...
- 4y ago
- typon 4y agoThere were several key insights that have made something like ChatGPT possible relative to traditional neural networks. * A fixed (but large) vocabulary of sub-word like tokens as inputs. * Attention mechanism for learning the correlation of words in a fixed sequence window. * Implementing this attention mechanism in the form of matrix multiplies rather than some other complex math - it allows it to be parallelized and run on GPUs fast. * Having enough layers of these encoders to have a huge amount of parameter space. ~175B parameters in the case of ChatGPT. * Feed the model a lot of data - in this case, pretty much the entire internet as text. * Self supervised learning: we take sentences from the internet and mask out some words and force the network to predict the missing word. Turns out this works extremely well. We don't use the traditional supervised learning inputs -> (prediction, label) paradigm that was the standard 10 years ago. * RLHF (Reinforcement learning from human feedback). Take generated text from GPT-3 (the underlying generative model) and ask humans to rate different completions. Retrain the model from those ratings. * A massive compute infrastructure that is able to train this model in a reasonable amount of time, allowing for iteration on hyperparameters. For example, what's the optimal attention head-size? How many encoder layers are good? What should the sequence length be? What should the embedding dimension be? etc. In OpenAI's case, they used thousands of GPUs and thousands of CPUs provided by Microsoft/Azure. In summary, relatively simple model, parallelizable on GPUs, trained on a lot of data.
- jupp0r 4y agoI asked ChatGPT and here is its answer: Sure, I'd be happy to help. An artificial neural network (ANN) is a type of machine learning algorithm that is designed to mimic the way the human brain processes information. At its core, an ANN is composed of a large number of interconnected processing nodes, called neurons, which are organized into multiple layers. These layers are interconnected in a way that allows them to pass information back and forth, with each layer building on the information provided by the previous layer to make a decision or generate a response. Transformers are a type of ANN that was introduced in 2017 by a team of researchers at Google. They are called "transformers" because they use a technique called "self-attention" to transform the input data before passing it through the network. This allows them to better capture the underlying relationships in the data, which makes them particularly well-suited for tasks like natural language processing and machine translation. ChatGPT is a type of transformer that has been specifically designed for use in chatbots and other conversational AI systems. It is an extension of the GPT-3 model, which was developed by the same team of researchers at OpenAI. GPT-3 is a large language model that has been trained on a massive amount of text data and is able to generate human-like responses to a wide range of questions. By building on the capabilities of GPT-3, ChatGPT is able to provide even more sophisticated and natural-sounding responses in a conversational setting. As for the size of the network, GPT-3 is a very large model, with 175 billion parameters. This means that it has a massive amount of interconnected neurons, which allows it to capture a wide range of knowledge and generate responses to a wide variety of questions. Of course, having a large network also requires a lot of computational power, which is why GPT-3 and other large language models are typically run on powerful servers or specialized hardware.
- DonHopkins 4y agoNow ask Ronald McDonald why his food is so healthy!
- chopete3 4y agoThis blog explains some of the key innovations they added on top of GPT-3, especially the natural language understanding (following instructions). https://openai.com/blog/instruction-following/ https://openai.com/blog/instruction-following/ In the first few paragraphs they show GPT-3 as equally dumb like all other language models that came before it and why they built instructgpt. >> Here is the summary (ChatGPT summarized) They present their approach to the problem, which involves using a recurrent neural network to encode both the instruction and the environment, and then using a reinforcement learning algorithm to learn how to execute the instruction. They demonstrate that their method can learn to solve a variety of instruction following tasks. Some snippets from the blog: >> InstructGPT is then further fine-tuned on a dataset labeled by human labelers. The labelers comprise a team of about 40 contractors whom we hired through Upwork and ScaleAI. >> We hired about 40 contractors, guided by their performance on a screening test meant to judge how well they could identify and respond to sensitive prompts, and their agreement rate with researchers on a labeling task with detailed instructions. We kept our team of contractors small because it's easier to have high-bandwidth communication with a smaller set of contractors who are doing the task full-time.
- chopete3 4y agoThe interesting thing was - there were 0 comments on HN when this blog was posted on Jan 27th. It was such a fundamental breakthrough NLU. NLU has been a such a holy grail of language models.
- stevenhuang 4y agoNLU?
- chopete3 4y agoNatural Language Understanding.
- deleted 4y ago[deleted]
- 4y ago
- birdyrooster 4y agoWhy didn’t you ask ChatGPT?
- wizofaus 4y agoRather strangely, it would seem - I just had this response: "However, I am a language model and do not have the ability to edit or revise my responses once they have been generated" Except I've had no problem getting it to do just that previously... I'm curious about its training data too, as I've managed to find a few things it knows nothing about (despite them having wikipedia pages and multiple dedicated websites about, and having been around 10+ years).
- BoorishBears 4y agoThey also trained a model on what qualifies as a "sensitive" prompt, which means: For a given prompt, the odds that it is considered sensitive is probabilistic Since it's an input outside if the prompting part, they can probably tune that specific aspect independently in response to controversial usages
- swayson 4y agoYannic Kilcher did an explainer recently on his YT channel. https://www.youtube.com/watch?v=0A8ljAkdFtg https://www.youtube.com/watch?v=0A8ljAkdFtg Yannic explains these models pretty well.
- karxxm 4y agoThis guy is awesome
- dotancohen 4y agoI just skimmed the video. It seems full of information, but this video skips. It seems like either the author composed his sentences by slicing together various clips, or a few frames have been removed between some words. I tried opening the video in NewPipe as well, same problem.
- rvbissell 4y agoI believe this is a stylistic, 'production values' choice. It used to bother me to the same degree that it seems to bother you. I suspect the content creators that do it are clipping out things like "umm" and "uhh". It might also be that some creators are better at this kind of splicing than others, which might explain why this particular example did not bother me so much.
- dotancohen 4y agoI see, thank you.
- touringa 4y agohttps://lifearchitect.ai/chatgpt/ https://lifearchitect.ai/chatgpt/
- timonoko 4y agoHow does the non-english languages part work? I thought maybe they use Google translator, but remembered that Russians have trained it to not to understand "russophobic" sentences. -- Mitä tarkoittaa ryssänvastainen, explain in English. -- Ryssänvastainen means "anti-Russian" or "anti-Russian sentiment." It refers to an attitude or behavior that is hostile or opposed to Russia or Russian interests.
- seydor 4y agoThe model says it has been trained in a variety of languages, but its current version is english. Its translations are bad, probably because it uses a translation model on top of its output. Sometimes it is obvious it is translating english articles from the mistakes it makes
- ravi-delia 4y agoThe original GPT-3 also did bad translations, just from scraps of foreign languages in the dataset. They seem to be worse though, so perhaps you're right that Chat-GPT is using another model the way it uses Internet searches
- int_19h 4y agoI don't know what it uses to output non-English text, but if you address it in a different language, it'll quite happily speak that back to you. And I'd say it understands it just fine, not really any worse than English - although it sometimes gets the word forms wrong in inflective languages.
- Scarblac 4y agoBut I was asking it for rhyming words in Dutch (because of Sinterklaas), and it suggested words that would rhyme if pronounced as English...
- bheadmaster 4y agoI also found the same thing happening in my native tongue. It would be much more exciting if it had "native" understanding of foreign languages.
- osigurdson 4y agoChatGPT is great. I use it a lot. But... it is still necessary to use google for many things. ChatGPT is somewhat out of date and strangest thing is, it will almost always provide an answer (right or wrong). So, for the most part, everything has to be manually checked.
- oars 4y agoHigh quality answers like in this thread are why I come to HN. Although I hope these high quality answers don't all just come from ChatGPT one day.
- layer8 4y agoFor better or worse, Ask HN is still wildly more useful than Ask ChatGPT.
- DemocracyFTW2 4y agoWhat baffles and makes it hard for me to wrap my brain around the statement that ChatGPT is "just predictive text analysis" is that it certainly feels like it goes into detail, picks up and repeats the particulars, presents a line of reasoning that sounds, well, reasonable, until you realize it got a number of things terribly wrong, often in deceptively subtle ways: Me: Which number is bigger, 873593742593454623 or 873593742593554623? ChatGPT: The second number, 873593742593554623, is larger than the first number, 873593742593454623. This is because the second number has a larger digit in the ones place (5 vs. 4), so it is greater overall. To compare two numbers, you can start by looking at the digits in the ones place and then move to the tens place, the hundreds place, and so on until you find a difference. The number with the larger digit in the place you are comparing is the larger number overall. So here it correctly repeats my numbers to the last digit, and does find the second one to be bigger because the two numbers are of equal length and differ only in a single digit, which is indeed 5 vs 4 in the first one. So far, it looks like ChatGPT really did an analysis rather than regurgitating old known phrases. However, that difference is not in the ones place, and the algorithm it gives is patently wrong (it doesn't mention that longer numbers are bigger than shorter numbers, and that numbers of equal length should be compared starting at the highest place, not the lowest). According to ChatGPT, 53 should be bigger than 123, so let's ask it: Me: Apply your reasoning for finding the bigger of two numbers to the compare 123 and 53. (Grammar mistake in the original) Me: To compare the numbers 123 and 53, we can start by looking at the digits in the ones place. In this case, both numbers have the same digit in the ones place (3 and 3), so we need to move on to the tens place. In the tens place, the number 123 has a larger digit (2 vs. 0), so 123 is the larger number overall. To compare two numbers, you can start by looking at the digits in the ones place and then move to the tens place, the hundreds place, and so on until you find a difference. The number with the larger digit in the place you are comparing is the larger number overall. This method can be applied to any two numbers, regardless of how many digits they have. That's somewhere between cringeworthy and deluded. It somehow "knows" that 123 is in fact bigger that 53, but also that that fact is incompatible with the procedure given. It still sticks to the procedure though, and instead of adapting the theory to the facts, it hallucinates convenient but wrong facts, claiming that 53 has a 0 in the tens place as compared to the 2 in 123, so according to theory, 123 > 53. Yeah I guess we should totally replace people with ChatGPT wherever possible.
- greenSunglass 4y agoHow much RAM/GPU/CPU is required to run the chatGPT / GPT3 model (aka text-davinci-003)?
- ilaksh 4y agoI'm using the text-davinci-003 model with the open AI API. It seems equivalent to chatGPT in terms of coding tasks. And the pricing is totally reasonable.
- infinityio 4y agoAs far as I know Davinci is supposedly a variant of gpt3-175B, so likely 175B parameters, so likely in the hundreds of gigabytes of ram unfortunately
- zeknife 4y agoIt is not available to run locally, but like the other comments said, an equivalent model would require an inordinate amount of VRAM.
- macrolime 4y agoIt used to be 2x number of parameters, so about 350GB. Facebook came with an optimization a couple months ago so it takes only half the memory, so about 175GB for GPT-3, not sure if OpenAI has implemented that yet. BLOOM, which is around the size of GPT-3, though not as good currently, works with these optimizations and should then be possible to run with 8x 24GB consumer GPUs. Huggin face has a nice blog post on this https://huggingface.co/blog/hf-bitsandbytes-integration https://huggingface.co/blog/hf-bitsandbytes-integration
- jb1991 4y agoChatGPT is trained using a combination of supervised and unsupervised learning. For supervised learning, it is trained on a large dataset of human-generated text, such as dialogue data or online conversations. This allows it to learn the structure and style of natural language. For unsupervised learning, it is trained using a language modeling objective, which involves predicting the next word in a sequence of text. This allows it to learn the broader patterns and characteristics of language, and to generate text that is fluent and coherent. ChatGPT and GPT-3 are both large language models trained by OpenAI, but they have some important differences. GPT-3 is a more general-purpose language model, which means it is trained on a broader range of data and can generate a wider range of responses. It is also much larger than ChatGPT, with 175 billion parameters compared to ChatGPT's 2.6 billion parameters. This makes GPT-3 more powerful and capable of generating more realistic and diverse text, but also makes it more expensive and resource-intensive to use. In case you are curious, the above information was written entirely by ChatGPT when asking it about itself.
- espadrine 4y agoThat is somewhat inaccurate. ChatGPT is based on the InstructGPT weights, based on the GPT-3 weights. It is roughly the same number of parameters, as far as we can tell. The GPT-3 weights were obtained by doing unsupervised pre-training (hence, GPT: generative pre-training): maximizing the likelihood of the model predicting the next word in a large dataset of human text. The InstructGPT weights were obtained with supervised fine-tuning (SFT) by making the model generate text, and asking a human to show a better text completion (as described in InstructGPT). Then, they also asked humans to rank multiple generated outputs, which was used as the supervised training goal of a separate reward function. That small amount of ranked data unlocked the ability to rank a much larger amount of data through reinforcement learning using proximal policy optimization (PPO): the model generates an output, the reward function rates it, and the model weights are updated to achieve a higher reward. The ChatGPT beta weights were obtained by doing that again, but asking the humans to make the completion conversational. Since they could only pay few humans, they opened the beta to ask a wider range of people to do SFT using the feedback feature, to train the final ChatGPT weights. So, all in all, the parameter estimation is incorrect, the order of the training is not right, the description of the purpose of the supervised learning step is wrong, the defining part of the ChatGPT training process is not mentioned (because the InstructGPT paper came in 2022, after the knowledge cut-off), the description of the difference with GPT-3 is misleading.
- ribit 4y agoI understand the basic idea of predicting the words in a sequence, but what totally eludes me is how this relates to the prompt. After all, you don't give it a sequence to continue, you give it a direct request. Is there some special processing going on here or do they really just take the prompt as is and encode it?
- Sharlin 4y agoA direct request is just a specific sort of a sequence to complete. There are extra training steps involved in ChatGPT that are designed to make it particularly good at predicting sequences that look like semantically valid, coherent dialogue, but fundamentally there's nothing special about dialogue. It's just tokens after tokens. For example, even a fairly simple model trained on only play and movie scripts would be naturally good at completing pieces of dialogue because that would be all that it knows!
- Trampoflix 4y agoPretty interesting talk about the foundation model used in chat GPT3. https://m.youtube.com/watch?v=D3sfOQzRDGM https://m.youtube.com/watch?v=D3sfOQzRDGM
- Trampoflix 4y agoHere is a pretty interesting talk about the foundation model used in ChatGPT. https://m.youtube.com/watch?v=r8ajJKDiT6s https://m.youtube.com/watch?v=r8ajJKDiT6s
- k__ 4y agoHalf-OT: people are always talking about ChatGPT being AI, but is this actually the case? It frequently told me that it doesn't learn from my input, and I had the impression the unique selling point of AI was it being able to modify it's own code in response to input.
- mjburgess 4y agoThat's basically never been the case for any statistical AI system which are all subject to "catestrophic forgetting". The weights of the system which are its memory are set by having their values determined over an entire training set. This is, effectively, a compression process which takes eg., most of the internet, and compresses it to 1TB of weights. If you update those weights by a process which doesn't "zip everything at once" then you're overwriting weight values by data which is less representative than "everything at once". I'd imagine in ChatGPT's case, OpenAI will review all the transcripts and find someway to do a pass over the whole set at once to make some improvement. If it were being trained live it's whole memory would be wiped out and replaced with racist/sexist/etc. inputs very quickly, as with MS's chatbot
- k__ 4y agoJust because it's a bad idea doesn't mean it's the right definition of AI. I always thought of the current systems of static machine learning models.
- Sharlin 4y agoThere is no exact definition of what "AI" is. But there have been millions and millions of software systems in the world that could classified as "AI" since the 60s, and approximately none of them have had any means of altering their code on the fly, so no.
- the_third_wave 4y agoIt does seem to learn during the session but it forgets what it has learned afterwards. I noticed this when I made a request in Dutch to which it replied in kind albeit with a grammatical error. I pointed out this error and showed it the correct form to which it replied by thanking me after which it consistently used the correct form during that session. When I reset the session and asked the same question it came back with more or less the same reply including the grammatical error.
- Doorstep2077 4y agoIt's definitely a step up from GPT-3, but I'm curious how much further it has to go before it's actually scary. Right now, I feel like there's still quite a bit of progress to be made.
- smallerfish 4y agoI like that as a measure of the AI field's progress. Impressive but not scary? We'd better keep developing!
- discordance 4y agoAnyone found a architecture diagram that includes the ML Ops parts? - I'm very interested in this at a system level for how the train / retrain loops work but haven't found much info on that.
- deleted 4y ago[deleted]
- chronolitus 4y agoBack when GPT-3 came out, I wanted to understand how it works, so read the papers and made this post: https://dugas.ch/artificial_curiosity/GPT_architecture.html https://dugas.ch/artificial_curiosity/GPT_architecture.html I hoped it would be simple enough for anyone who knows a bit of math / algebra to understand. But note that it doesn't go into the difference between GPT-3 and ChatGPT (which adds a RL training objective, among other things).