4 ms·
I've only seen people mention that it runs really slow, even on like A100s.
by bestcoder69 3y ago
I've only seen people mention that it runs really slow, even on like A100s.
- microtonal 3y agoThere are no big differences compared to other LLM architecturally. The largest differences compared to NeoX are: no biases in linear layers, shared heads for the key and value representations (but not query). Of course, it has 40B parameters, but there is also a 7B parameter version. The primary issue is that the current upstream version (on Huggingface) hasn't implemented key-value caching correctly. KV caching is needed to bring the complexity down from O(n^3) to O(n^2). The issues are: (1) their implementation uses Torch' scaled dot-product attention, which uses incorrect causal masks when the query/key sizes are not the same (which it the case when generating with a cache). (2) They don't index the rotary embeddings correctly when using key-value cache, so the rotary embedding of the first token is used for all generated tokens. Together, this causes the model to output garbage and it only works when using it without KV caching, making it very slow. However, this is not a property of the model and they will probably fix this soon. E.g. the transformer library that we are currently developing supports Falcon with key-value caching and it the speed is on-par with other models of the same size: https://github.com/explosion/curated-transformers/blob/main/curated_transformers/models/refined_web_model/layer.py https://github.com/explosion/curated-transformers/blob/main/... (This is a correct implementation of the decoder layer.)
- zwaps 3y agoSuper helpful! Do you have some more info on these issues and where this is discussed? Besides following y'all at explosion, any tipps of whom to follow so I don't get blindsided?
- microtonal 3y agoI found them out myself when making our own implementation of the model. We test our outputs against upstream models. In decoding without history, our tests passed, but in decoding with history there was a mismatch between our implementation and the upstream implementation. Naturally, I assumed that our implementation was wrong (being the newer implementation, not sharing code with theirs), but while debugging this I found that our implementation is actually correct. Then I was planning to report these issues. Someone else found the causal mask issue a week earlier, so there was no need to report it: https://github.com/pytorch/pytorch/issues/103082 https://github.com/pytorch/pytorch/issues/103082 I reported the issue with rotary embeddings in a discussion of problems that people were running into trying to use KV caching: https://huggingface.co/tiiuae/falcon-40b/discussions/48#648c828eb8f4a3542b7abd4a https://huggingface.co/tiiuae/falcon-40b/discussions/48#648c... More in general, I am not sure what the best place is to track these issues. Maybe a model's discussion forums?
- kbrkbr 3y agoI tried it using oobabooga's webui side by side with Alpaca 65B loaded in 4 bit on the same AWS instance with 64GB of VRAM. While Alpaca produced 3 tokens/sec, Falcon produced 0.17 tokens/sec. So it is very slow with the current tooling still.
- brianjking 3y agoHow can you deploy Oogabooga to AWS/Huggingface/etc? Any tips? Cheers!