3 ms·
Inference should ideally be done on an instance with a GPU, and it should have enough GPU VRAM to hold the entire model. This one is only a 7b model, so it shou
by cheald 3y ago
Inference should ideally be done on an instance with a GPU, and it should have enough GPU VRAM to hold the entire model. This one is only a 7b model, so it should run pretty easily in a modest amount of vram. You can likely run it locally in CPU-only mode, too, though it's likely to be rather slow.
https://instances.vantage.sh/?min_gpus=1 https://instances.vantage.sh/?min_gpus=1 (wait for it to finish loading to filter)
You could try running these in Google Collab (though tweaks would have to be made to load the files), or you might try something like runpod.io, as well, which gets you GPU instances for a lot less than you'd pay with AWS - for example, a Tesla V100 (16GB) community cloud instance runs you $0.24/hr vs the g5g.xlarge which runs you $0.42/hr.