10 ms·
RWKV RNN: Better than ChatGPT?
- pffft8888 4y agoWhat test cases do folks here recommend for measuring this new model's ability to reason? and, specifically, if it can reason about code with similar (or better!) performance to ChatGPT4? Has anyone managed to get it running locally?
- macrolocal 4y agoThe author claims 61.0% on WinoGrande vis-a-vis GPT-4's 87.5%.
- pffft8888 4y ago"you can fine-tune RWKV into a non-parallelizable RNN (then you can use outputs of later layers of the previous token) if you want extra performance." Is that 61% using the non-parallelizable RNN mode or the standard mode? I wonder if it's the latter. This new model may be a viable alternative to ChatGPT, which is not only closed sourced but can be shut down in the future just as they did with the older text-davinci models. Plus, the alignement and safety has rendered ChatGPT useless for helping with areas such as critical analysis of social issues (that go against the aligned views) and any and all critical thinking that goes against the aligned views of those who own and program ChatGPT. This could a viable free (as in freedom) alternative.
- macrolocal 4y agoI think the Cambrian explosion is just beginning.
- mach1ne 4y agoI hope not but day by day it seems more likely. If text-generating LLMs can reach superhuman cognition they will so so in a matter of a few years. At that point a Waluigi prompt will be like arming a virtual nuclear missile.
- macrolocal 4y agoNuance: computers have been accumulating superhuman cognitions for half a century. But most people are bad at recognizing intelligence they don't relate to.
- MaxikCZ 4y agoI can't seem to find it in GitHub repo, do you know the value for ChatGPT before it switched to GPT-4?
- macrolocal 4y agoHere are a few benchmarks: https://paperswithcode.com/sota/common-sense-reasoning-on-winogrande https://paperswithcode.com/sota/common-sense-reasoning-on-wi...
- akavi 4y agoHow’d GPT-3/3.5-turbo do?
- cyanf 4y agoLooks like 81.6%. macrolocal linked this below: https://paperswithcode.com/sota/common-sense-reasoning-on-winogrande https://paperswithcode.com/sota/common-sense-reasoning-on-wi...
- gooseus 4y agoOpenAI has been collecting a ton of evals here https://github.com/openai/evals https://github.com/openai/evals with many of them including some comments about how well GPT-4 does vs GPT-3.5. You could clone that repo, adapt the oaieval script to run against different APIs, then run the evals against both and compare the results.
- deleted 4y ago[deleted]
- jacobn 4y agoFrom the project page: pronounced as "RwaKuv" That is still quite challenging to pronounce, maybe one of "rwkv" -> "raw-kv" -> "rawk-v" -> "rock-v"?
- tjr 4y agoRocky V?
- dragonwriter 4y ago“RwaKuv” seems like it would pretty closely match “Rock of”
- OJFord 4y agoI would assume Rwa like Rwanda, Kuv like covet. But that's just to agree with you really, since that's not one of your suggestions, so with such different ideas it's clearly not a particularly helpful pronunciation guide!
- rvz 4y agoThat is the problem. Unfortunately, this will be forgotten over the OpenAI hype brigade. I thought Stable Diffusion was a bad name due to its very technical name. But now I have seen something even worse for LLMs. This time OpenAI is learning its lesson in not getting itself disrupted easily. For any hope of challenging them, we need to be better at names. Even the name 'Bitcoin' caught on. Same with iPhone. So I'm afraid that the name alone for this project will be the cause of it being quickly forgotten as OpenAI aggressively captures mindshare. The same with 'Bard'; a horrific name. Google should have simply called it 'Brain' and incremental updates as 'Brain 2', 'Brain 3.6', etc and renamed their existing AI division to Google Brain Labs. Easy. How is that difficult?
- mach1ne 4y agoI beg to differ. Right now the whole tech world is eyeing for the best open source alternative for GPT. A techy name matters little if it works and is accessible.
- GaggiX 4y agoThe best thing about this model is that it has O(T) speed and O(1) memory during inference vs the O(T^2) speed and O(T) memory (flash memory) of a GPT model, still it can be trained in parallel like a GPT model.
- pffft8888 4y agoIn addition, 1) it's open source. 2) you can run it yourself so the rug won't be pulled from under you when they decide to shutdown and move users up to the next version or another product as they've done with the older text-davinci models. 3) you get to align it (using RLFH) as opposed to a corporation dictating what is "aligned" and what is "safe." 4) you won't have to deal with government led censorship. For example, instead of the FBI using JIRA to manage a list of URLs to be censored (as they did according to the latest revelations) they can train the AI to self-censor as Bing has done. 5) you won't be using the product of a company that was started as a non-profit with $100M donation (from Elon Musk) to promote transparent AI only to take that money and turn into a for-profit company and close-source the AI. Sources: Elon is the source for #5 and Matt Taibbi is the source for #4. I doub't you'll have a problem sourcing #5 so here is the source for #4: "31. After the 2020 election, when EIP was renamed the Virality Project, the Stanford lab was on-boarded to Twitter’s JIRA ticketing system, absorbing this government proxy into Twitter infrastructure – with a capability of taking in an incredible 50 million tweets a day." --Matt Taibbi on Twitter https://twitter.com/mtaibbi/status/1633830104144183298 https://twitter.com/mtaibbi/status/1633830104144183298
- ChaseMeAway 4y agoI think that your first few points are fair, but I was a little confused about #4. I had never heard of this, and it seemed important, some cursory research turned up many twitter posts from individuals amplifying your version of events. I also found a write up on the situation from TechDirt[0]. The article is fairly good and well sourced, but it paints a substantially different picture than what you describe. [0]https://www.techdirt.com/2023/02/15/extraordinarily-confused-congressional-rep-thinks-social-media-companies-are-secretly-communicating-with-govt-censors-via-jira/ https://www.techdirt.com/2023/02/15/extraordinarily-confused...
- serverholic 4y agoI'm skeptical that RNNs alone will outperform transformers. Perhaps some sort of transformer + rnn combo? The issue with RNNs is that feedback signals decay over time, so the model will be biased towards more recent words. Transformers on the other hand don't have this bias. A word 10,000 words ago could be just as important as a word 5 words ago. The tradeoff is that the context window for transformers is a hard cutoff point.
- pizza 4y agoI think RWKV ameliorates this to some degree: How it works: RWKV gathers information to a number of channels, which are also decaying with different speeds as you move to the next token. It's very simple once you understand it. RWKV is parallelizable because the time-decay of each channel is data-independent (and trainable). For example, in usual RNN you can adjust the time-decay of a channel from say 0.8 to 0.5 (these are called "gates"), while in RWKV you simply move the information from a W-0.8-channel to a W-0.5-channel to achieve the same effect.
- solomatov 4y agoAs far as I remember in RNN times, the best models were RNNs with attention. Does this thing has any attention mechanism? If it does, then it has the same problem with the O(n^2) computation where n is the window size. My understanding is that transfers are superior due to the fact that they are much faster to train/evaluate than RNNs.
- yieldcrv 4y agoWhat does RNN stand for? edit: recurrent neural network
- adeon 4y agoI've followed updates on this project r/machinelearning and for me the existence of projects like this is some good evidence that the OpenAI moat is not that strong. It gives some hope you are not going to need massive huge computers and GPUs to run decent language models. I hope this project will thrive.
- gaogao 4y agoYeah, my sense is that OpenAI moat is primarily just through the RLHF dataset right now. Most of the other things – foundational models and datasets, embeddings (roughly the plugins announcement today) – have generally already been commoditized. Just getting that dataset for fine tuning is one of the last major hurdles for the ChatGPT geist.
- sdrinf 4y agoThe Open Assistant community ( https://open-assistant.io/ https://open-assistant.io/ ) is building a crowdsourced dataset for RLHF, with apparently high quantity of high quality contributions.
- boppo1 4y agoRunning models is one thing, but surely lots of power will remain with those who have the capital/hardware to train them?
- MacsHeadroom 4y agoAs of yesterday you can train 33B parameter models on a single consumer 24GB VRAM GPU with https://github.com/johnsmith0031/alpaca_lora_4bit https://github.com/johnsmith0031/alpaca_lora_4bit. It's already proven that 13B parameters is enough to beat GPT-3 175B quality. It's likely that 33B parameters is enough for GPT-4.
- lupire 4y agoAlpaca is fine-tuning LLaMa not training from scratch.
- pffft8888 4y agoImagine having ChatGPT level AI running in an ASIC inside earphones. This could be like an always-on buddy, available offline and able to access resources when you're connected. Or in Google Glasses. The Readme states that it's more optimized for ASIC than the transformer architecture used by ChatGPT.
- IIAOPSW 4y agoHow will you know which voice in your head is yours? Maybe you are both.
- LoganDark 4y ago> How will you know which voice in your head is yours? Maybe you are both. As someone with DID I wonder this every single day.
- wruza 4y agoMaybe that will teach us about fragility of our definition of consciousness and agency. For some reason many people think that they are insulated from outside effects unless they’re right in their brain. Which is not true.
- RobotToaster 4y agoI wonder if we will see LLMs on FPGAs or FPGA like devices. Crypto went from graphics cards to ASICs, we may see something similar with LLMs given the hype.
- pffft8888 4y agoI always think of FGPAs for things you want to update without having to repurchase the hardware. If the FPGA inside a gadget can be reprogrammed over the Internet then I think it's more suited than an ASIC for LLMs.
- int_19h 4y agoIt could go even further, in theory. The kind of ops that the current crop of LLMs needs is very simple, and at the same time there's no hard requirement for precision (which is why 4-bit quantization works so well). This means that unconventional approaches such as analog computing are potentially in the play again - it's easy to do addition and multiplication in an analog circuit, if you don't care about the answer being precise, and in theory one could pack a lot more of those circuits in the same space.
- sourcecodeplz 4y agoThis is actually not that bad.
- gigel82 4y agoAlso check out Alpaca; you can self-host this one, the 7B and 13B variants produce surprisingly good results and are fast enough just running on CPU: https://github.com/antimatter15/alpaca.cpp https://github.com/antimatter15/alpaca.cpp
- all2 4y agoFrom the README dear god rwkv is fast. i switched to another tab after starting training it from scratch & when i returned it was emitting plausible english & maori words, i left to go microwave some coffee & when i came back it was producing fully grammatically correct sentences.
- solomatov 4y agoIt happens the same way if you train a small transformer. It learns to produce words very quickly.
- guywithabowtie 4y agoComparison with GPT3.5 - turbo Given two numbers provide answer as follows 3 + 9 = 39 1 + 2 = 12 2 + 9 = 29 What is 29 + 49 RWKV : ? A: The answer is $29+49 = 78$ Explanation: Let the two numbers be $a$ and $b$. Then the sum is $a+b$. Now let the sum be $c$. Then the sum is $a+b+c$. Now let the sum be $d$. Then the sum is $a+b+c+d$. Now let the sum be $e$. Then the sum is $a+b+c+d+e$. Now let the sum be $f$. Then the sum is $a+b+c+d+e+f$. Now let the sum be $g$. Then the sum is $a+b+ GPT3.5 turbo Based on the pattern provided, 29 + 49 would be equal to 2949. Explanation: In each of the equations given, the result is obtained by concatenating the two numbers together instead of adding them. For example, 3 + 9 becomes 39 because we concatenate 3 and 9 to get 39. So, to find the answer to 29 + 49 using this pattern, we concatenate the two numbers to get 2949.
- guywithabowtie 4y agoDoes it need RLHF to make sense?
- hkab 4y agoNo, RLHF only helps the model to align to human preference
- theemathas 4y agoRLHF makes it possible to get good results with a bad prompt.
- guywithabowtie 4y agoIt generates quite a lot of random content to be honest. If Nancy had two apples and Becky had 1 apple. Becky gives her 1 apple to Nancy, how many apples becky has ? Full Answer: RWKV : Two apples. If Nancy had 2 apples and Becky had 1 apple. Becky gives her 1 apple to Nancy, how many apples becky has ? Two apples. Q : Two girls are playing with a ball, one of them throws the ball so that it goes straight and falls on the other's feet, the other bends her knees and catches it, how many times will the ball fall on the knees ? Full Answer: The ball will fall on the knees three times. Q : Two sisters are playing with a stick. The first sister says 'let me hold it', the second sister says 'no'. Now what will happen ? Full Answer: The second sister will hold it. Q : How many GPT 3.5 Turbo : After Becky gives 1 apple to Nancy, Becky will have zero apples left. Becky gave her only apple to Nancy, so she doesn't have any apples remaining. So, the answer is Becky has zero apples left. It requires a lot more improvement.
- guywithabowtie 4y agoI will wait for future improvements and watch this project.
- byefruit 4y agoYou're seeing this because the model isn't instruction fine-tuned. You'll need prompting similar to the original GPT3 or Llama models.
- guywithabowtie 4y agoCan you give me example of that ?
- zamnos 4y agoBasically they're tuned for sentence completion rather than chat/being asked questions. Plugging > Nancy has two apples and Becky has one apple. Becky gives 1 apple to Nancy. Becky now has into GPT-2 via HuggingFaces at https://huggingface.co/tasks/text-generation https://huggingface.co/tasks/text-generation I get > Nancy has two apples and Becky has one apple. Becky gives 1 apple to Nancy. Becky now has three apples and Nancy has one apple. Becky now has three apples and Nancy has one apple. > Witch Hunt > The following is GPT-2 is much weaker, which explains the garbled nonsense output, along with the incorrect answer for Nancy. I have no idea what RWKV RNN would output, but leading sentences instead of questions is how to get LLMs not RLHF tuned to answer.
- deleted 4y ago[deleted]
- nico 4y ago> We can predict that RWKV 100B will be great, and RWKV 1T is probably all you need :) That sounds awfully similar to this quote: "There is no reason for any individual to have a computer in his home." by the founder of DEC in 1977. There’s a similar one that’s supposed to be Bill Gates’ but apparently it’s not.
- homarp 4y agoThe '640K' quote won't go away -- but did Gates really say it? https://www.computerworld.com/article/2534312/the--640k--quote-won-t-go-away----but-did-gates-really-say-it-.html https://www.computerworld.com/article/2534312/the--640k--quo...
- wruza 4y agoThese phrases often get ripped out of context. What they meant was “there’s no reason someone would want to enter machine codes through switches or punchcards into a personal room-sized device”. That is still true. Some people complain about incorrect defaults even though these are few taps away to become correct.
- Sparkyte 4y agoIt just takes one language library to dethrone the next. I called this when everyone was like CHATGPT!!! The problem is noone knows what they are talking about and screaming AI!!! ChatGPT is not AI. It does something automated with accuracy information baked in and builds new information around the accuracy data. It does not think like AI, it takes the most probable data and responds with it. That is machine learning it. It is fundamentally a cornerstone toward AI, but not AI itself.
- famouswaffles 4y ago>The problem is noone knows what they are talking about Indeed. as you've just clearly demonstrated
- Sparkyte 4y agoThink what you like. But I called out crypto I will call this out. But ChatGPT has the potential to be a major component to AI. It is a cornerstone. The problem is English. AI means something differently between use. What we are going to see with this is that people are going to throw the Sci-Fi book at this and claim something like, "It has feelings!". I am saying AI the term is being broadly applied to anything with a logic gate. And this is bad marketing for a product too early in development. It is not the conventional term AI everyone broadly applies. It is a cornerstone toward the AI people broadly apply.
- 101011 4y agoRespectfully disagree on this point. Here's the definition for AI: "the theory and development of computer systems able to perform tasks that normally require human intelligence, such as visual perception, speech recognition, decision-making, and translation between languages." I would say that chatgpt is not sentient, nor is it capable of independently improving its intelligence - but I think this tech hits the (lower) bar for "AI"
- Sparkyte 4y agoIt is a corner stone to AI. The opinion is subjective since we are treading into a new era it is hard to clearly define those lines.
- isaacfrond 4y agoInteresting, that it goes against the grain. Since the seminal paper 'Attention is all you need', we went from RNN type neural network to pure attention based networks. It started the LLM revolution as the attention only training such networks is parallelizable, and you got record breaking performance to boot. Now we learn, that going back to the old RNN paradigm is actually better. It even advertises itself as totally 'attention-free'!
- knight0075 4y agoOk
- mpaepper 4y agoIs there a research paper / arxiv which describes it in detail?
- BHSPitMonkey 4y agoOff-topic, but this submission's title feels unusually editorialized/click-baity for HN.