13 ms·
The genie escapes: Stanford copies the ChatGPT AI for less than $600
- twblalock 4y agoThis is why it's not possible to slow down or "stop" AI: once the problems are solved the solutions turn out to be trivial to replicate. All it takes is compute.
- jeron 4y agoyou say all it takes compute like that is trivial - chatGPT would have a hard time without Microsoft's support via Azure
- crooked-v 4y agoWhile that's true, it's basically inevitable now that at some point personal hardware will be powerful enough for enthusiasts to run home bots comparable to GPT-3, and even that by itself would drastically change a lot things.
- layer8 4y agoRunning isn’t necessarily the issue. The moat is creating a high-quality model like OpenAI has, which (and here the article is mistaken) doesn’t seem to be easily reproducible.
- crooked-v 4y agoWhile that's true, it also seems entirely predictable at this point how to do that. It takes a lot of effort and expensive hardware, but there isn't really a "secret sauce" beyond expertise in the field.
- layer8 4y agoYes, but it takes time (took OpenAI years) and significant effort. Who with enough expertise will do this and not keep the results closed in order to monetize them? It doesn’t seem like something an open source project could accomplish quickly enough to not keep lagging substantially behind the commercial solutions.
- twblalock 4y agoThat's going to get easier too. Stanford can already get this far for $600, so soon after the major GPT-based chat AIs were released. Imagine how much better it will get with just a little bit more time.
- drowsspa 4y agoSounds like IBM trusting no one would copy their BIOS code.
- bmacho 4y agoGovernments can ban powerful devices, as they can ban guns, bombs, and such.
- AnimalMuppet 4y agoOnly at a price. If you ban devices powerful enough to run ChatGPT, you ban a big chunk of what powers your economy.
- zamnos 4y agoMaybe. Certainly in the past, before the world was aware LLMs on the level of ChatGPT were possible with today's technology. OpenAI's chosen not to release any real details about GPT-4, so we don't actually know what it would take to train a model of equivalent quality, especially considering training isn't a one-shot. Multiple training runs easily add up training costs. So training for a 12-figure parameter size model(s) (175B) is assumed to be very expensive. But there has been great progress made for optimized models which are smaller by a two orders of magnitude - 7B for a debatable drop in quality (7B alpaca is in no-way competitive with ChatGPT, but it's still very much not a markov chain from during the AI winter). So one possibility is that OpenAI chose not to release salient GPT-4 details is due to it being much smaller than GPT-3's 175B model size and they're hiding the details because of how much that cuts down on training costs. (Which I should note is unsubstantiated conjecture but not outside the realm of possibility.) The other aspect is that fine-tuning an existing model is way cheaper than creating a competing model from scratch, so a company could offer CompetitorGPT/CompetitorCoPilot competitive with GPT-3.5, and offer fine-tuning of that model trained on the source code repository of the purchaser company's codebase, possibly on-prem or at least inside their AWS VPC/Azure/GCP equivalent. The other thing to note is that OpenAI is hosting ChatGPT as a public resource available to anyone with an account, akin to Google being open to the public from day one (although that is without an account. Maybe Gmail is a better comparison). I can't say for certain, only OpenAI would know for sure, but I'm willing to bet that inference for ChatGPT is the vast majority of their costs (which is all but trivial). Any private internal-only instance of OpenChatGPT (using the unlicensed leaked LLaMA model or a legal copy or someone else's) could be paying (relatively) minuscule training costs, and way lower inference costs if it's internal-use only. Whether that cost can be borne by a small SaaS company's existing AWS budget is up in the air, which is to say ultimately that you're right - ChatGPT would be difficult without the support of Microsoft via a huge Azure grant, it's less obvious that a self hosted internal-only OpenChatGPT, not from OpenAI, would be possible by hobbyist self-hosters with a prosumer GPU cluster (Say with last generation K80's instead of business-priced A100's), or by a company wanting to leverage LLMs for private use by that company that wants to provide a Copilot like productivity multiplier internal tool to their developers, without sending private source code to OpenAI in lieu of a privacy agreement with them.
- 4y ago
- twblalock 4y agoThere are lots of places to get compute, including Chinese cloud providers... The genie really is out of the bottle now. This is a lot like pharmaceuticals. The initial investment in a new medication is enormous. The price of each pill is trivial, to the extent that every drugstore chain is able to supply a generic in-house brand.
- mLuby 4y agoGovernments have experience limiting the spread of digital content. For now at least, AI proliferation is not immune to those same tactics.
- superkuh 4y agoHardly. I've played a lot with the 7,13, and 30B llamas as well as the 7 and 13B alpacas fine tuned by Stanford. They do not have emergent abilities like being able to generate rhymes or, say, represent a movie plot as emoji. Even openai's old text-davinci-003 (gpt3.5, but text completion, not the chat ones) far outperforms them. That said, I have hopes for a 65B 3-bit quantized alpaca-fine tuned. We'll see when someone spends the money to do the (more costly) 65B training. The alpacas are also much more likely to go off rails and start regurgitating their fine-tuning inputs. Either that or openai is doing a lot of post processing on their end to hide the same problems in their LLM. For now my IRC bots run the alpaca 7B 4-bit. 13B was not a significant improvement for twice the computational time. But it's best to learn them now because as soon as openai gets sued for the first time all the turing test passing older models without the legal-butt-covering bolted on will be removed.
- nickthegreek 4y agoWhere does one find the 13B alpaca model?
- superkuh 4y agoBe aware this file is a single ~8GB 4-bit model (ggml-alpaca-13b-q4.bin) instead of the 2x ~4GB models (ggml-model-q4_0.bin, ggml-model-q4_0.bin.1) that most llama.cpp style inference running programs expect. You'll probably have to edit the line, n_parts = LLAMA_N_PARTS.at(hparams.n_embd); in chat.cpp (or main.cpp) to hard code it to treat this 1 file model properly like, n_parts = 1; Or re-write the parameter config subroutine to recognize and handle non-standard weights file. magnet: magnet:?xt=urn:btih:053b3d54d2e77ff020ebddf51dad681f2a651071&dn=ggml-alpaca-13b-q4.bin&tr=udp%3A%2F%2Ftracker.opentrackr.org%3A1337%2Fannounce&tr=udp%3A%2F%2Fopentracker.i2p.rocks%3A6969%2Fannounce&tr=udp%3A%2F%2Ftracker.openbittorrent.com%3A6969%2Fannounce&tr=udp%3A%2F%2F9.rarbg.com%3A2810%2Fannounce torrent: https://btcache.me/torrent/053B3D54D2E77FF020EBDDF51DAD681F2A651071 https://btcache.me/torrent/053B3D54D2E77FF020EBDDF51DAD681F2... torrent: https://torrage.info/torrent.php?h=053b3d54d2e77ff020ebddf51dad681f2a651071 https://torrage.info/torrent.php?h=053b3d54d2e77ff020ebddf51... via: https://github.com/antimatter15/alpaca.cpp https://github.com/antimatter15/alpaca.cpp
- starik36 4y agoFrom the article: Pre-trained on a trillion "tokens"... Doesn't 7B indicates that it was trained on 7 billion tokens? Or am I misunderstanding the nomenclature?
- instance 4y ago7B is the number of parameters of the model.
- superkuh 4y agoThe emerging consensus for larger LLM is you want to train them with at least 2-4x the tokens of the number of parameters (weights between neurons in the layers). A trillion (100x) surprises me.
- sebzim4500 4y agoThey probably put most of the effort into the 65B model, the 7B model was just trained so they could get an idea of the scaling behaviour. It makes sense to use the same amount of training steps, then.
- sitic 4y agoThe LLaMA paper contradicts this view: "[...] Although Hoffmann et al. (2022) recommends training a 10B model on 200B tokens, we find that the performance of a 7B model continues to improve even after 1T tokens." https://arxiv.org/pdf/2302.13971.pdf https://arxiv.org/pdf/2302.13971.pdf
- dragonwriter 4y ago> Doesn't 7B indicates that it was trained on 7 billion tokens? No, 7B means it has 7 billion parameters.
- starik36 4y agoAnd what does a parameter mean in this context?
- raydiatian 4y ago> It seems these godlike AIs are already frighteningly cheap and easy to replicate. Who writes this shit?
- B1FF_PSUVM 4y agoHard to say ...
- dr_kiszonka 4y agoThanks for a genuinely funny comment!
- meh8881 4y agoIrreplaceable humans
- welly34h 4y agoCode can be abstracted into a simpler code model and deterministically recreate the old code model. OpenAI is an eventually to be obsoleted initial brute force approach that will be abstracted over and over into a simpler code implementation with rules to recreate the old state. kkrieger is a simple example of a tiny data model that can be deterministically rehydrated. It’s not unrealistic for AI models to become a seed value for a normalized code base to deterministically unpack into necessary electron state
- freediver 4y agoThe incredible contribution of Alpaca is showing the world how to efficicently train LLM on instructions. The fact that it did so on 52k instructions generated by GPT is poetic. It does not matter what current capabilities of open source models are, because this opens the door to tremendous democratization of the ability to train and self-deploy these models. In less than 6 months we will have open source models with gpt3-like capabilities, running locally on laptops, and potentially in phones and web browsers.
- deleted 4y ago[deleted]
- permo-w 4y agoif we’re all still alive by then
- braingenious 4y agoI have not found alpaca to be comparable to chatgpt, but it could be because of bugs in the version I installed through dalai. I might try reinstalling it because I suspect there might be some sort of file corruption issue or whatever. I gave it the prompt “cats aren’t always fuzzy” and it wrote a lengthy livejournal-esque rambling journal entry about a woman and her husband having money issues. It was funny, but lightyears away from chatgpt. It does sometimes create some really funny hallucinations though, like inventing prefectures in Japan that don’t exist etc.
- DustinBrett 4y agoI also got that text about the married couple and their money issues. Alpaca didn't impress me at all so far.
- disgruntledphd2 4y agoAlpaca wasn't great. The 13b and 30b models are much better, but just for sentence completion. Personally, I think that the RLHF does make a big difference but maybe it's a bug in the quantization code as suggested up thread.
- braingenious 4y agoI’m also a bit confused by the quantization thing. Why exactly is everybody running the same program on the same file? Why not just include the quantized weights? It seems like if somebody figured out the “correct” way to quantize the 7b weights it would make way more sense to just torrent the output rather than distribute a fixed program.
- homarp 4y agoquantization takes lots of RAM: https://github.com/qwopqwop200/GPTQ-for-LLaMa/blob/main/README.md https://github.com/qwopqwop200/GPTQ-for-LLaMa/blob/main/READ... says llama-13B takes 42GB and 33B takes more than 64GB...
- awinter-py 4y ago> asked GPT to take 175 human-written instruction/output pairs, and start generating more in the same style and format ... through one of OpenAI's helpfully provided APIs, and ... the team had some 52,000 sample conversations to use in post-training the LLaMA model hmm I wonder if this is essentially a probe[1] technique + relies on chatgpt already having been extensively trained like did they basically exfiltrate the weights 1. probing per https://arxiv.org/abs/2102.12452 https://arxiv.org/abs/2102.12452
- doctoboggan 4y agoI've used both the 7B and 13B instruction tuned llama weights (quantized using the llama.cpp scripts). Either I am doing something wrong, or these two models are no-where near the level of ChatGPT. Many times they return something totally irrelevant to my question, stop responding, use a different language, or otherwise return the wrong answer. ChatGPT does none of this. (other than the wrong answer due to hallucinating sometimes...) Reading through the README and issues on the llama.cpp project, there is some speculation that there is a bug in the quantization, or possibly a bug in the inference (less likely I think). I hope this is true and once fixed the models can perform up to or past the ChatGPT level. If its not true and these models are performing correctly, then either the metrics used to compare it to GPT is garbage and don't capture the real world uses, or the instruction tuning done by the Stanford team is not up to par.
- f_devd 4y agoLLama hasn't been fine-tuned with RLHF, so it requires additional prompting, check out the open-assistant[0] project for an open-source ChatGPT equivalent (WIP). [0]: https://github.com/LAION-AI/Open-Assistant https://github.com/LAION-AI/Open-Assistant
- simonw 4y agoThis is why Alpaca is a big deal: it shows what LLaMA can do after it's been fine-tuned to follow instructions like ChatGPT has.
- f_devd 4y agoAlpaca uses Self-Instruct[0] which is better than just the pre-training but I wouldn't expect it to be at the level of ChatGPT (RLHF) in terms of human-friendly prompting. OpenAssistant should make it close to ChatGPT (from GPT-3.5 version) if the LLaMA is as powerful as claimed. [0]: https://arxiv.org/abs/2212.10560 https://arxiv.org/abs/2212.10560
- satvikpendem 4y agoUse it here: https://huggingface.co/spaces/olivierdehaene/chat-llm-streaming https://huggingface.co/spaces/olivierdehaene/chat-llm-stream...
- earthboundkid 4y agoI bet you could “exfiltrate” an LLM relatively cheaply by using LLM A to generate training data for LLM B.
- oezi 4y agoNo way. The cost for generating the tokens is way too high.
- jakedata 4y agoAI bootstrapping AI is a sci-fi trope that goes back decades. I first encountered it in The Cybernetic Samurai while in high school. While the details differ, the reality is that AI is a catalyst for more of itself. I don't remember many books where this ends particularly well. Perhaps the Culture universe could be a survivable outcome. Hopefully we don't get Berzerkers first.
- nuclearsugar 4y agoStable Diffusion trains StyleGAN2 - https://www.jasonfletcher.info/vjloops/ https://www.jasonfletcher.info/vjloops/
- doctor_eval 4y agoIt’s the beginning of the AI singularity. It’s not that it’s bad, we just can’t see anything beyond the event horizon.
- EGreg 4y agoI warned about this for years. Finally an article gets it right. Everyone will soon have the equivalent of online nuclear weapons: bot swarms that infiltrate every forum, including this one.
- tantalor 4y agoSpam has existed on the internet for a long time.
- EGreg 4y agoThis is different. It can act just like humans do for most people who skim comments won't be able to tell the difference. Note this was in 2020: https://www.technologyreview.com/2020/10/08/1009845/a-gpt-3-bot-posted-comments-on-reddit-for-a-week-and-no-one-noticed/ https://www.technologyreview.com/2020/10/08/1009845/a-gpt-3-... And here's 4chan bot: https://www.youtube.com/watch?v=efPrtcLdcdM https://www.youtube.com/watch?v=efPrtcLdcdM I can tell you that HN is probably already being infiltrated as well. SPAM can't gang up on you in a forum and downvote you and turn your friends against you and destroy your reputation within 1 hour online. But soon, it will. The web as we know it is soon going to be over.
- tantalor 4y agoHow good are these bots against CAPTCHA?
- B1FF_PSUVM 4y agoMy pet theory was that AI would come out of spam bots. Close enough.
- Waterluvian 4y agoIf you use consciousness as a baseline, the intellectual difference between a grade schooler and a PhD is tiny. This is what I think comparing these bots is like. You can argue that they’re very close. But the delta makes a very big difference for any practical purposes because we’re looking for nuanced capability.
- est31 4y agoThere is a blog post that drives this point home with very good illustrations: https://waitbutwhy.com/2015/01/artificial-intelligence-revolution-1.html https://waitbutwhy.com/2015/01/artificial-intelligence-revol... https://waitbutwhy.com/2015/01/artificial-intelligence-revolution-2.html https://waitbutwhy.com/2015/01/artificial-intelligence-revol... Basically, at the point where we have "almost human" level AI, it won't take much to get AI that's beyond human capabilities.
- cjohnson318 4y ago> It seems these godlike AIs are already frighteningly cheap and easy to replicate. "godlike"? Really? I'm not religious, but this seems like an overreaction for something that has no agency.
- Jcowell 4y agoWhat if creation was a result of a lucky happen-by-chance hallucination ?
- cjohnson318 4y agoSure! Why not? Quantum foam imagining quantum foam. I love it. I still wouldn't consider an LLM god-like. I mean, if it eats its own son, then maybe. (Titan of a reference there.)
- crazygringo 4y agoIf it's a shorthand for omniscience then I can see how it makes sense. A bit hyperbolic though for sure.
- doctor_eval 4y agoHow do you know agency is not simply the output of a large language model encoded in neurons? What is the difference between neuronal and digital weights? Considering that we don’t know how the brain works so well, and we don’t understand why LLMs work so well, simply on the basis of their output I think the safest assumption is that these models do indeed have agency, or at least the capability of agency.
- cjohnson318 4y agoI'm using agency in this context to mean: (1) a strong desire to achieve one or more clear goals, beyond survival, and (2) taking concrete steps to achieve those goals. > How do you know agency is not simply the output of a large language model encoded in neurons? I'm not sure what you mean here. Is agency an emergent effect of large digital or biological neural network? Maybe! Is it an emergent effect of a large language model? If it is, then it should be clear, or demonstrable, that the model (1) has goals (2) takes concrete steps to achieve those goals. > What is the difference between neuronal and digital weights? Brain chemistry works at orders of magnitude less speed, since we're talking about periodically building and releasing an ionic differential between the inside and outside of a cell wall. Moreover, we have a massive number of neurons and a stupidly massive amount of interneuronal connections, with billions of years of training over billions of lineages. Digital weights, in contrast, are a stripped down model of this system that throws out a whole class of complexities like hormones and metabolism. > I think the safest assumption is that these models do indeed have agency, or at least the capability of agency. I think this is an overly generous assumption.
- UncleOxidant 4y agoIs it accurate to say they were trained for less than $600? Wouldn't that just be the finetuning that was done to the already existing LLaMA parameters which likely cost way more than $600 to train?
- simonw 4y agoYeah, exactly. LLaMA 7B itself cost $80,000+ to train (82,432 GPU hours). Stanford spent $100 on fine-tuning compute and $500 on OpenAI credits to generate their 52,000 sample instruction training set.
- simonw 4y agoRelated, my post "Could you train a ChatGPT-beating model for $85,000 and run it in a browser?" https://simonwillison.net/2023/Mar/17/beat-chatgpt-in-a-browser/ https://simonwillison.net/2023/Mar/17/beat-chatgpt-in-a-brow... I think you can train LLaMA 7B (the model underlying Alpaca) for around $82,000, based on the Meta Research paper about it. Then you can fine-tune it ala Alpaca for a few hundred dollars more. My wilder speculation is that, if you can shrink the model down to 4GB with llama.cpp 4bit quantization, it may be possible to run it entirely in the browser (ala Stable Diffusion from the other day).
- dang 4y agoRecent and related: Stanford Alpaca web demo suspended “until further notice” - https://news.ycombinator.com/item?id=35200557 https://news.ycombinator.com/item?id=35200557 - March 2023 (77 comments) Stanford Alpaca, and the acceleration of on-device LLM development - https://news.ycombinator.com/item?id=35141531 https://news.ycombinator.com/item?id=35141531 - March 2023 (66 comments) Alpaca: An Instruct Tuned LLaMA 7B – Responses on par with txt-DaVinci-3 - https://news.ycombinator.com/item?id=35139450 https://news.ycombinator.com/item?id=35139450 - March 2023 (11 comments) Alpaca: A strong open-source instruction-following model - https://news.ycombinator.com/item?id=35136624 https://news.ycombinator.com/item?id=35136624 - March 2023 (296 comments)
- satvikpendem 4y agoAlpaca is cool but it's also not technically allowed by OpenAI's TOS, and LLaMA is certainly not allowed to be used for non-commercial purposes. With that in mind, OpenAssistant is an Apache 2.0 licensed fully open source alternative that's pretty good (the model is OpenAssistant/oasst-sft-1-pythia-12b): https://huggingface.co/spaces/olivierdehaene/chat-llm-streaming https://huggingface.co/spaces/olivierdehaene/chat-llm-stream.... I've found OA to be better than Alpaca but I'll wait until the 65B 3-bit quantization efforts for Alpaca are underway to compare them.
- Zuiii 4y ago> Alpaca is cool but it's also not technically allowed by OpenAI's TOS, and LLaMA is certainly not allowed to be used for non-commercial purposes. Only if you agreed to the ToS or believe that the weights are copyrightable (precedents set by the copyright office and the courts strongly suggest that they aren't). I personally see no issue in using these models for commercial purposes.
- satvikpendem 4y agoYou might not but a company will think twice. It's the same reason why companies could theoretically use pirated Windows and Adobe products and get away with it, but most don't because the risk is not worth the reward.
- Zuiii 4y ago> It's the same reason why companies could theoretically use pirated Windows and Adobe products Again, using and distributing LLaMA weights is not illegal in any way under current laws. End of story.
- satvikpendem 4y agoAgain, you missed the point. LLaMA is under a non commercial license. I never stated it is illegal, just that a company will not want to violate license terms even if it is legal, simply because getting sued in a civil case is still a risk they wouldn't be willing to take compared to the benefit.
- neilellis 4y agoWow, Stanford's Alpaca AI project is a real game-changer. The fact that it performs on par with ChatGPT but costs less than $600 to build is both exciting and terrifying. Sure, it's great to see AI becoming more accessible, but it's also a massive wakeup call for the potential misuse of these technologies. We've got big names like OpenAI, Google, Apple, Meta, Baidu, and Amazon putting in serious time and money to ensure their language models are safe and ethical. However, now that we know it's possible to build powerful AI models on a budget, it's crucial to think about what this means for the future of AI regulation and safety. This Alpaca AI project is a stark reminder that we need to have a serious conversation about the possible repercussions of AI proliferation. We can't just sit back and assume the big companies will take care of everything. The genie is out of the bottle, and it's time for everyone in the tech community to face the music and take responsibility for the AI revolution.
- amrb 4y agoAnything open will need training and attention to be an openai competitor, tho I'm happy to see the function of this one: https://huggingface.co/spaces/togethercomputer/OpenChatKit https://huggingface.co/spaces/togethercomputer/OpenChatKit
- xwdv 4y agoGiven the high prices of OpenAI offerings it seems it’s better to pirate an AI model before resorting to paying for anything.
- ronsor 4y agoChatGPT API is surprisingly cheap, but GPT-4 is many times more expensive to the point where I can't see it being worth it most of the time.
- stavros 4y agoI'm not convinced that the GPT-4 API actually works? It's been giving me very different answers than the chat interface. For example, the chat interface says it's GPT-4 if you ask it, but the API says it's GPT-3 (and bills as GPT-4).
- siva7 4y agoIf you're satisfied with the results who cares what it thinks it is
- stavros 4y agoI'm not. It's worse than the chat version.
- johntb86 4y agoWhy would the LLM model know what version it actually is? It was trained on data from before that model version even existed.
- stavros 4y ago- What version model are you? - I am based on the GPT-4 architecture, which is an iteration of OpenAI's Generative Pre-trained Transformer models. My training data includes information up to September 2021. - How do you know your version? - As an AI language model, I am programmed with information about my own architecture and the version I am based on. My creators at OpenAI provided me with this knowledge, which helps me to give context to users about my capabilities and limitations. In this case, I am based on the GPT-4 architecture, an iteration of OpenAI's Generative Pre-trained Transformer models.
- gaogao 4y agoHas anyone tried this yet on the 65B version? I'm curious if it knows how to rhyme and other emergent behavior, as alpace-7B does not.
- starik36 4y agoI tried because I was curious. It's not very good. First I "quantized" the files. Took about 3-4 hours. Then I ran a prompt: "create a short poem about my wife helping me get over the stuffed nose". After 3 hours of 99% CPU usage and using all 32 GB of RAM, I got the following result: create a short poem about my wife helping me get over the stuffed nose and throat that i've had for days now. the cold was all around him, but he didnΓÇÖt notice it at first- so much to do in this new place! finally though when they were sitting by themselves on their porch swing his sweetheart leaned into her man saying ΓÇ£we could take some of these leaves from the garden and make tea.ΓÇ¥ He looked up with a quizzical look and saw that she was serious. I've been under weather since last Tuesday. Today is day 6...and still going strong! Had to cancel two nights on stage, missed my son https://i.imgur.com/Nl3xLEg.png https://i.imgur.com/Nl3xLEg.png
- alecco 4y agohttps://archive.ph/xIKIN https://archive.ph/xIKIN