3 ms·
How much time does this take to get produced locally?
by cced 3y ago
How much time does this take to get produced locally?
- smoldesu 3y agoI run a similar Vicuna bot on a free ARM VPS from Oracle. Inference takes ~20 seconds for ~200 tokens on the CPU, so it should stream results about half as fast as ChatGPT. ...however, that was on cheapo ARM hardware. On a hospital budget I bet you could beat OpenAI with a local model pre-mapped in GPU memory on a 3090.
- cced 3y agoDo you guys have any reference material on how to get this started?
- smoldesu 3y agoJust one guy, me :) But there are a lot of different routes to go. If you want to do what I've done though - 1. Sign up for Oracle Cloud and try to get their free 4 core ARM VPS (these are at-capacity very often, but free if you can get them) 2. Install the Ampere Pytorch runtime: https://cloudmarketplace.oracle.com/marketplace/en_US/adf.task-flow?tabName=O&adf.tfDoc=%2FWEB-INF%2Ftaskflow%2Fadhtf.xml&application_id=125935163&adf.tfId=adhtf https://cloudmarketplace.oracle.com/marketplace/en_US/adf.ta... (can also use the ONNX or Transformers acceleration image if you need it) 3. Install your Pytorch/ONNX/Transformers software, link in system libraries and enjoy! Whole thing works super smooth in my experience. 7B and 13B models are very usable for chatbot type applications.
- amstan 3y agohttps://github.com/lm-sys/FastChat https://github.com/lm-sys/FastChat The hardest part is downloading the 30GB of weights.
- thatcherc 3y agoNot OP, and I'd have to see the prompt to run it on my set up, but I but together a refurbished HP Chromebox with some extra RAM and an SSD for a total of $250 and it'll run all the 13B models about as fast as I can type on my phone, just as a reference point. I'm consistently impressed by it even if it's not blazing fast.
- cced 3y agoDo you guys have any reference material on how to get this started?