7 ms·
This is super awesome, too bad a A6000 cost $4500 or I would try this out myself.
by syntaxing 4y ago
This is super awesome, too bad a A6000 cost $4500 or I would try this out myself.
- moyix 4y agoThe smaller models run on smaller GPUs too! :) You can see how much VRAM is required for various models in the Documentation: https://github.com/moyix/fauxpilot/blob/main/documentation/server.md#setup https://github.com/moyix/fauxpilot/blob/main/documentation/s... And we have an initial GPU support matrix page here: https://github.com/moyix/fauxpilot/wiki/GPU-Support-Matrix https://github.com/moyix/fauxpilot/wiki/GPU-Support-Matrix A planned feature is to implement INT8 and INT4 support for the CodeGen models, which would let the models run with much less VRAM (~2x smaller for INT8 and ~4x for INT4) :)
- Namidairo 4y ago> A planned feature is to implement INT8 and INT4 support for the CodeGen models, which would let the models run with much less VRAM (~2x smaller for INT8 and ~4x for INT4) :) That would certainly put the larger models within reach of most users. Part of my reluctance to deploy this myself was I'd have to either redline the VRAM on a 8GB card or choose the smallest model. Would this require a retrain of the models?
- moyix 4y agoIt shouldn't require retraining, nope. I believe for INT4 there is a small adapter layer added that needs to be trained, but it is small and wouldn't require much data or computation to do so. Will know more once we actually start implementing :)
- daviding 4y agoThe 7GB runs great on a 3080Ti. I am getting a lot of 'ValueError: Max tokens + prompt length' errors with larger files. Can this Gitlab client also replace the vocab.bpe and tokenizer.json config like Copilots? Thanks for your work on Fauxpilot, really enjoying playing with it.
- moyix 4y agoI believe right now the VSCode extension just passes along the entire file up to your cursor [1] rather than trying to figure out how much will fit into the context limit – it's definitely still very early stages :) It would be pretty simple to run the contents through the tokenizer using e.g. this JS lib that wraps Huggingface Tokenizers [2] and then keep only the last (2048-requested_tokens) tokens in the prompt. If they don't get to it first I may try to throw this together soon. [1] https://gitlab.com/gitlab-org/gitlab-vscode-extension/-/blob/main/src/completion/gitlab_code_completion_provider.ts#L57 https://gitlab.com/gitlab-org/gitlab-vscode-extension/-/blob... [2] https://www.npmjs.com/package/tokenizers https://www.npmjs.com/package/tokenizers
- daviding 4y agoUnderstood - thanks again. (plus note to self; have individual source files with less in them ;)
- KaoruAoiShiho 4y agoAhh, 3090/4090 conspicuously missing, those affordable 24gbs are very common in the ML community, please add them to the matrix.
- moyix 4y agoI don't have any to test on unfortunately! But it's a wiki; if you get some of the models running on those (it should be as easy as running ./setup.sh) please add a line saying that it works!
- teruakohatu 4y agoTwo 3090 GPUs could probably handle it, making the cost slightly cheaper.
- KaoruAoiShiho 4y agoIs it possible to do 2x 4090s or is 3090s the only option putting 2 together?
- moyix 4y agoYes, you can do 2x4090s as well. NVLink is not required (though it will make things a bit faster).
- KaoruAoiShiho 4y agoSweet great to hear.
- biggerChris 4y ago[flagged]