Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
pico_creator
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
61.
▲
by
pico_creator
3y ago
I strongly believe. Given the resources to scale RWKV to the size of LLaMA and beyond The performance advantages without compromises will be so glaringly obvious (over 10x) that the industry will change forever ——— Ps: I am bias, I am the s
62.
▲
by
pico_creator
3y ago
Ahh. As someone who has been listed as “independent researcher” before. I get what you mean - oh well, it is what it is
63.
▲
by
pico_creator
3y ago
I would not really call it random. While it was open for feedback/contributions. There is a strong requirements for substantial contribution to the paper itself to qualify for authorship. So unfortunately that does limit it in part to
64.
▲
by
pico_creator
3y ago
wierdly their form requires a company rep (which RWKV does not have, as its not a company) - lets see how it goes ...
65.
▲
by
pico_creator
3y ago
The paper is meant to compare architecture vs architecture with similar model size, and dataset - to inform decisions in future architecture designs Its main benefits being presented with the above staying closely the same, is that it has s
66.
▲
by
pico_creator
3y ago
Yup, if you follow this definition of attention. It makes sense. The mixing step, is computed with the trained weights, meaning the model does learn on its own, when/what to emphasise in this "mixing" process. Hypothetically
67.
▲
by
pico_creator
3y ago
If you have more specific feedback, like a specific digram or page, and how it can be made better. I will gladly forward that info, to improve the paper draft. Because channel mixing, is a core component of this architecture, and that keywo
68.
▲
by
pico_creator
3y ago
weirdly enough, organisations are more willing to rent GPUs than money. If you want to help fund RWKV, the ko-fi link is - https://ko-fi.com/rwkv_lm IMO: this needs way more funding, just to sustain blink leading this proje
69.
▲
by
pico_creator
3y ago
Haha, yea - naming things is hard Every-time someone comes up and say "this is not attention, because it does X and not Y" My response is, ok, what would you call it then? Because no one (including me) seem to to able to find a be
70.
▲
by
pico_creator
3y ago
Instruction training. This is a WIP
71.
▲
by
pico_creator
3y ago
For anything past 8k context size We are talking about over 10x reduction in GPU time for inferencing tokens and for training too Aka it’s cheaper and faster Alignment is frankly IMO purely a dataset design and training issue. And has nothi
72.
▲
by
pico_creator
3y ago
Really bad napkin math as no one has attempted 65B (so +\- 50%) 8 x 8 x 8 A100, should be able to do a 100k++ tokens/s at that size With a dataset of 1.2 trillion tokens. That’s 12 million seconds. Or 140 days (PS: this is why everyone
73.
▲
by
pico_creator
3y ago
Prompt design is definitely a huge one shifting to rwkv Sadly too many folks copy and paste what works for openAI and move on when it fails
74.
▲
by
pico_creator
3y ago
Hmm we might need to look into the instruct training data. Which is mostly based on gpt4all filtered and mixed with others (You are using raven right? That’s the instruct trained varient) Btw ping the discord if ur looking into finetuning f
75.
▲
by
pico_creator
3y ago
You mean red pajama? I believe that has already started for 1-14B (need to double check)
76.
▲
by
pico_creator
3y ago
Agreed. IMO - A part of me even argue we should stop calling it attention (but what to call it instead is a mess) But since this was derived from apple lite attention paper. The name is gonna stick, due to a lack of better alternative
77.
▲
by
pico_creator
3y ago
Chinchilla law is a rule of thumb that you should have 11++ x training tokens for every param If not, you are getting diminishing benefits for each param you add I’m extreme cases your model can even perform worse with more param due to lac
78.
▲
by
pico_creator
3y ago
It does not exists (yet maybe) Blink is an individual and does not represent a company (aka not google, not eleuther, not <insert VC company>, etc) So he had to fill something up i guess haha The idea of a foundation has been tossed a
79.
▲
by
pico_creator
3y ago
Yup. This is commonly done in the community for the chat models as well (due to the huge amount of reuse for each reply)
80.
▲
by
pico_creator
3y ago
Hahaha. True. And even then it’s an undersell (the limited scope of the paper drops lots of the things being done)
81.
▲
by
pico_creator
3y ago
Rearrange the query. Ask the question / explain the task first. Then give it the data you want to extract from. Also you may want to give a one shot example for best result (the instruct training of raven model is very limited)
82.
▲
by
pico_creator
3y ago
Dun we all wish this wasn’t the case? Where we have more OSS models to choose from without weird rule lawyering gotchas. Or needing to be from a research institute / a license to download the weights
83.
▲
by
pico_creator
3y ago
That’s a loaded question without deciding dataset size
84.
▲
by
pico_creator
3y ago
Taking into account the limits of how much can be stored in a hidden state. My theory is that the model will start generalising the information stored once it starts going past its limits (like real life humans) So if we take books as an ex
85.
▲
by
pico_creator
3y ago
If you are familiar with how transformer network works There is RWKV in 150 lines to help understand all the nitty gritty https://github.com/BlinkDL/ChatRWKV/blob/main/RWKV_in_150_li...
86.
▲
by
pico_creator
3y ago
TLDR: please donate A100s to make this happen Most of the focus is in the 1-14B range. Due to constraints of the dataset sizes (chinchilla law), and GPUs available Community demand is also mostly in this range as there is a strong desire to
87.
▲
by
pico_creator
3y ago
A large percent of the RWKV community ain’t experts. And are here doing weird, dumb or crazy homebrew experiments So keep doing weird experiments
88.
▲
by
pico_creator
3y ago
Layperson summary Good - substantially cheaper and faster to run / train - scalable to ridiculous context size Bad - you will need to change how you prompt this model sadly (it works differently)
89.
▲
by
pico_creator
3y ago
Yea using the analogy of the eyes can see the entire document Vs I need to memorise everything as it’s spoken out, and then answer on it Is a good approximate on the difference. The kicker though as people pointed out, there are individuals
90.
▲
by
pico_creator
3y ago
There are lots of really low hanging fruits - integrating this with AI platform X/Y/Z - setting up evals - improving the code quality - making a how to guide (it’s stuck on my todo list) - helping with dataset - doing silly experi
More ›