4 ms·
Has anyone tried this on an m1 machine?
by lachlan_gray 4y ago
Has anyone tried this on an m1 machine?
- tmptmptmp1 4y agoDo you have access to the weights? If so you probably have better ML hardware. Wish this model was actually "open". The perf of a model of that size on the M1 will not be good. That is big enough it won't quite fit on a 3090 (24GB) without quantization.
- pumanoir 4y agoI think is feasible. The description even says is designed to save on vram[1]. I don't get the other comments about needing more vram than a 3090. Also, Neuralmagic may run their sparsification on ARM cpu's in the future, so keep an eye. 1. ChatRWKV v2: with "stream" and "split" strategies. 3G VRAM is enough to run RWKV 14B :)
- IanCal 4y agoYou have to split it up which slows it down a lot. The 14B model doesn't fit fully on a 3090, though the 7B fits easily and is very fast. Other replies either may have meant this or thought the original comment was about llama.
- fswd 4y agoI've ran it on a AMD 3950 which I think is half the speed of a M1, and it's plenty fast. Note I am specifically talked about RWKV
- nl 4y agoIt would be interesting to see a version of RWKV[1] that takes some of the improvements in LLaMA (eg the SwiGLU activation function and the Rotary Embeddings - although I think they have tried rotary embeddings in some versions of RWKV) as well as the same dataset and see how it does. The dataset is interesting. It's not dissimllar to The Pile, which RWKV is already trained on, but does seem to have quite a lot more preprocessing to increase the dataset quality. [1] https://github.com/BlinkDL/RWKV-LM https://github.com/BlinkDL/RWKV-LM