16 ms·
Open source solution replicates ChatGPT training process
- raydiatian 4y ago> “the generative-AI eruption” I really think we should stick to Nick Bostrom’s (or pls fix attribution) term “intelligence explosion”
- SunghoYahng 4y agoEven if it has not so much thing to do with intelligence?
- raydiatian 4y agoI’m not sure about your definition of intelligence. Perhaps you think I’m saying ChatGPT and generative agents are somehow conscious. I don’t conflate consciousness with intelligence here. I can’t say whether or not ChatGPT is conscious (although I doubt it), but it’s pretty clearly intelligent by a reasonable definition. It’s an agent which is extremely effective at playing its game. A game which is incredibly open ended from the human point of view, whether or not the resulting agent’s internal model is based on statistical patterns. Consciousness is not a prerequisite to intelligence. But back to what I’m really saying here: “Generative AI eruption” is a mouthful whereas “intelligence explosion” is concise.
- plutonorm 4y agocongrats on the level headed response to "is it conscious". We do not know is the only correct answer. I'm getting very annoyed with many simply stating that they aren't conscious without any real understanding. Folk scientism.
- mach1ne 4y agoJust to play the devil's advocate here, you can argue that ChatGPT is not really intelligent. It is one hugely complex probability distribution, true, but importantly, a static one. Human intelligence stretches the similar distribution into the temporal dimension as our brain processing data influences the shape of the distribution real-time. In this manner, human brains are not Chinese Rooms due to their continuous online learning, but ChatGPT is, since the weights could be written to stone and the output would not be any different.
- raydiatian 4y agoI don’t think I buy that this is playing the devil’s advocate, or that it’s even a meaningful argument. 1) Whether the weights are static or dynamic over time is not of importance. As a simple counter argument, if for instance the theory that LLMs could produce AGI if they could just be pushed to an absolutely colossal scale, then a planet scale computer might produce a machine intelligence by the definition of this conversation. That’s a big what if, and it’s about as useful as string theory, but it illustrates something I touch on in point #2. 2) A second counter argument to the “well the weights are hardcoded and ChatGPT doesn’t learn” argument, ChatGPT does learn. I’ve taught it conversational protocols, where it stored information mutates it, and lets me retrieve it in a standard I invented on the spot. This is the entire basis of ChatGPT, understanding call-response of human conversation in the probabilistic abstract. The apparatus ChatGPT uses to “stretch the similar distribution into the temporal dimension” is that it can store new information passively in the continuing conversation thread. You could theoretically teach ChatGPT about 2023 by having a conversation about recent events. It probably wouldn’t be as effective as having trained it on new information, but nonetheless. 3) Finally, a deeper argument against what you’re saying, what you’re arguing has nothing to do with the definition of intelligence I’m using here. That is: “An agent which is exceedingly effective at its game.” it’s important to characterize intelligence by the game that the agent in question aims to play. When we say a child is intelligent, we don’t mean it in the same way that we say a doctor is intelligent. The two are playing entirely different games, yet both are considered intelligent. This is because the game parameterizes intelligence, and the intelligence explosion is all about the proliferation of specialized intelligences for the variety of “games” out there. ChatGPT is exceptional at its game of “general human conversational prediction up to the year 2021”.
- 93po 4y agoI'd argue computers from the early 1900s or even before are intelligent. Intelligence is pretty broad. Human intelligence is a very specific type of intelligence.
- rvz 4y agoFinally, an open-source equivalent to ChatGPT emerging out of the AI euphoria will begin to extinguish the hype out of OpenAI's ChatGPT moat, just like how GPT-3 and DALLE-2 were almost immediately disrupted by open-source models as well. This (and other open-source AI models), not 'ChatGPT', 'DALLE-2', etc is what will change the AI landscape for everyone, permanently forever.
- supriyo-biswas 4y agoI, for one, would like to see an open-source model similar to Stable Diffusion, but for text. It would be a great way to empower general folk without having to pay OpenAI, and not have to worry about the LLM's belief system, which is conservative-biased in the case of ChatGPT[1] (HN discussion[2]). [1] https://davidrozado.substack.com/p/openaicms https://davidrozado.substack.com/p/openaicms [2] https://news.ycombinator.com/item?id=34625001 https://news.ycombinator.com/item?id=34625001
- return_to_monke 4y agothere is https://github.com/laion-ai/open-assistant https://github.com/laion-ai/open-assistant being built in the open already. you can contribute too. please also notice that the article you linked is about the text classifier of the frontend and not the LLM itself
- mach1ne 4y agoThat's what I love about this particular AI revolution. The technologies are developed in such a non-siloed manner that open source is able to replicate the largest steps forward in a manner of a year.
- return_to_monke 4y agoto be fair they really are not there yet. They are just in the "data collection" phase, the actual training and then tuning is still to do. but hey, those are the same people who made the dataset (laion5b) for stable diffusion. I have hope.
- VadimPR 4y agoHow good is the quality of this? BLOOM is a 176B parameter model, but it doesn't seem to compare to GPT-3 (175B parameters) in terms of output quality.
- lossolo 4y agoIt's because BLOOM is undertrained, you can prune a lot of weights in BLOOM and it doesn't impact performance. Look at Chinchilla paper[1], 70B model outperforms 175B GPT-3 model. https://arxiv.org/abs/2203.15556 https://arxiv.org/abs/2203.15556
- Der_Einzige 4y agoIn general, most giant LLMs are extremely undertrained at this time. Consider that most of the gains in RoBerta vs bert were from just continuing to train.
- stevenhuang 4y agoCases of undertraining can be observed whenever the output is repeating gibberish or loops. Happened a lot in GPT2 ai dungeon days
- leobg 4y agoSo can we continue training RoBERTa to get it to, say, GPT3 Ada level
- rnosov 4y agoOut of curiosity, how did your measure their respective performances? My understanding is that BLOOM roughly comparable to GPT-3 in performance on most NLP tasks. Were you comparing OpenAI davinci to raw BLOOM by any chance?
- VadimPR 4y agoCompared ChatGPT to BLOOM - which I know doesn't benefit from RLHF.
- simonw 4y agoIs the term "ChatGPT" being used in place of GPT-3 here? Is this thing actually replicating the GPT-3 training process? The thing that makes ChatGPT interesting (over regular GPT-3) is the RLHF process, but this article doesn't seem to touch on that at all, unless I've missed something.
- de6u99er 4y agoGPT-3 has been publicly covered in scientific publications. Same as GPT-2, and GPT. Those are all pre-trained models, where GPT is the abbreviation of Generative Pretrained Transformer. Transformers have been invented in 2017 at Google Brain [1]. -> https://medium.com/walmartglobaltech/the-journey-of-open-ai-gpt-models-32d95b7b7fb2 https://medium.com/walmartglobaltech/the-journey-of-open-ai-... GPT-4 is around the corner, and it's allegedly 100x more powerful than it'd predecessor. -> https://medium.com/geekculture/gpt-4-100x-more-powerful-than-gpt-3-38c57f51e4e3 https://medium.com/geekculture/gpt-4-100x-more-powerful-than... [1] https://arxiv.org/abs/1706.03762 https://arxiv.org/abs/1706.03762
- wcoenen 4y agoThat source about GPT-4 is nonsense. It claims GPT-4 will have trillions of parameter, and at the same time links to another page which says that it won't be much bigger than GPT-3: https://www.datacamp.com/blog/what-we-know-gpt4 https://www.datacamp.com/blog/what-we-know-gpt4
- simonw 4y agoThat "100x" figure is extremely poorly sourced. I don't believe that at all.
- de6u99er 4y agoYou're right. Apologies for that.
- imtringued 4y agoAnd yet the intimidating pictures of a small and large circle keep getting posted everywhere.
- simonw 4y ago"hitting 100 million monthly active users 2 months after its launch". I'm deeply suspicious of that number. It came from Similarweb, who track these things through analytics gathered from browser extensions. I trust this article more: https://www.nytimes.com/2023/02/03/technology/chatgpt-openai-artificial-intelligence.html https://www.nytimes.com/2023/02/03/technology/chatgpt-openai... "But two months after its debut, ChatGPT has more than 30 million users and gets roughly five million visits a day, two people with knowledge of the figures said." "Two people with knowledge of the figures" is journalism speak for "I heard this off the record from people with insider info, and I'm ready to report it because those two different sources provided the same number".
- jackblemming 4y agoCan someone tell me what the hell they use ChatGPT for? I tried it a few times and it always confidently gave me wrong results to basic things. What is this thing supposedly “disrupting”? Is it really just marketing cranking out metric tons of spam blogs?
- savolai 4y agoI'm using it for crud, i.e. generating insert sql from c++ classes. Knows how to do acid compliance it seems with multiple tables and foreign keys, saving lots of time. It's also the better english to finnish translation than gtranslarw. Also copywriting as certain genres are highly repetitive.
- simonw 4y agoSo many things. A lot of them for personal entertainment, but increasingly for useful other stuff too. I used it to help brainstorm talk titles and abstracts for a talk I was proposing the other day. What I ended up submitting was entirely written by me but was heavily influenced by the ChatGPT conversations. https://til.simonwillison.net/macos/sips https://til.simonwillison.net/macos/sips - I used it to figure out how to convert webp to PNG on macOS, and learned about an entirely new built-in command. I often use it as a thesaurus - "what's a good word / term for X?" I'm self-employed and a journalist asked me for my job title, which I don't have. So I brainstormed some ideas with ChatGPT. I pasted in the output of a SQLite "explain query plan" query and asked for an explanation - which helped me figure out enough to write a section of this TIL: https://til.simonwillison.net/sqlite/subqueries-in-select https://til.simonwillison.net/sqlite/subqueries-in-select This is just from the past few days.
- jacooper 4y agoIm not deep into the AI space, but who would I use this? Do I just run it and speak to it in terminal? Or what is the next step to make it useful for search or more?
- rnosov 4y agoTFA claims that they managed to replicate "RLHF" type thing that would allow you to bark orders at the raw GPT-3 model and get palatable results back (as opposed to often repetitive nonsense output of the raw model). You won't be able to run this in your terminal as GPT-3 alone consumes nearly 400GB of RAM plus whatever post processing you do with it. At the moment, there is no obvious use case for it apart from running a ChatGPT competitor. On the other hand, there was no obvious use case for electricity for nearly a century. But one can speculate that we're getting closer and closer to a "lightbulb" moment for AI.
- college_physics 4y agoIt would somehow be combined with an open source search engine
- gnramires 4y agoI wish I could @marginalia_nu here :)
- sillysaurusx 4y ago> On a single multi-GPUs server, even with the highest-end A100 80GB GPU, PyTorch can only launch ChatGPT based on small models like GPT-L (774M), due to the complexity and memory fragmentation of ChatGPT. Hence, multi-GPUs parallel scaling to 4 or 8 GPUs with PyTorch's DistributedDataParallel (DDP) results in limited performance gains. Where are these numbers coming from? An 80GB A100 GPU is certainly more than capable of hosting a 1.5B GPT. We were running 774M on rinky-dink cards back in 2019 for our inference purposes. I don’t understand how they went from talking about 175B params across 32 cards to 774M on one card. 175B divided by 32 is 5.4B. In fact, I’m not sure what they’re saying in general. They seem to be confusing data parallelism with model parallelism with memory fragmentation, while namedropping a bunch of training techniques. The hard part of ChatGPT isn’t the size. It’s the training process. It took a small army of contractors rating outputs as good or bad. Once that dataset gets replicated, we can start talking about size. Hopefully LAION will deliver.
- sdenton4 4y agoYeah.... Having spent a lot of cycles replicating ML work, it's much more difficult than taking a stab at replicating a paper. It's typically doable (results really do replicate) but it can take a few good brains a year to pull it off. There's typically a lot of small decisions that add up, and a lot of hyperparameter sweeps to land in a good region of the optimization space.
- popinman322 4y ago> Once that dataset gets replicated, we can start talking about size. Hopefully LAION will deliver. Is LAION starting a community project to rate model outputs? I didn't see anything on their site.
- sitic 4y agoHere it is: https://open-assistant.io https://open-assistant.io (https://projects.laion.ai/Open-Assistant/ https://projects.laion.ai/Open-Assistant/)
- rnosov 4y ago
- LoganDark 4y agoIs cool~ Waiting for the day when I can run a model like this in a native language like Rust, without incurring the overhead of the Python interpreter. Python can be good for trying out methodologies but it's sort of yucky to set up ime.
- charcircuit 4y agoThe python interpreter is not the bottleneck
- AbusiveHNAdmin 4y agoThe graph titled "Comparison between Colossal-AI and current major open source projects in the same period" has no label on the Y axis, which shows quantities in thousands. WTH?
- rnosov 4y agoY axis are Github stars. They sort of mention it in the preceding paragraph.
- AbusiveHNAdmin 4y agoThanks for the reply.
- college_physics 4y agoWhy are gazillions of parameters needed in the first place? From an information perspective it feels that there might be some fundamentally inefficient use of parametric freedom. A brute force approach to combinatorial explosion so to speak. Are there any research efforts that look into how to reduce model complexity (without substantially sacrificing performance obviously).
- rnosov 4y agoOne way to think about it is that the model needs to essentially encode the entirety of human knowledge. If you can do it with just 175b parameters then it looks quite efficient to me. GPT-3 is about 400gb in size which would even fit in some modern IPhones! Another metric to consider is that there are about 100 trillion connections in the human brain. If you roughly equate brain connection to a model parameter then GPT-3 would be only 0.175% size of human brain.
- college_physics 4y agoA model parameter is not the same as a "fact". Facts can multiply uncontrollably, but the logical relationships between facts that (at least we as humans) care about are much more economical. It feels that this approach is missing some key abstractions that might help reduce redundancy in encoding. But its just a hunch. Need to dig deeper to understand at least conceptually why this dimensional explosion.
- imtringued 4y agoGiven a large enough model, model architecture becomes increasingly less relevant as any specialized architecture can be discovered by the larger model automatically. The only benefit of a specialized architecture is minimizing resource usage.