5 ms·
That already exists depending on your definition of slow. Just get a big ssd, use it as swap and run the model on cpu.
by chessgecko 4y ago
That already exists depending on your definition of slow. Just get a big ssd, use it as swap and run the model on cpu.
- leereeves 4y agoA comment below said this model uses fp16 (half-precision). If so, it won't easily run on CPU because PyTorch doesn't have good support for fp16 on CPU.
- netr0ute 4y agoParent never claimed it was going to be fast.
- leereeves 4y agoIt would probably just fail with an error "[some function] not implemented for 'Half'"
- chessgecko 4y agofp16 models inference just fine in fp32, though I was sorta joking in my original comment, it would potentially take weeks for this to run one input. You're better off trying to make something like huggingface accelerate work (like the comment above), which swaps layers of the model on and off the disk