4 ms·
How do you test a 70B model locally? I've tried to query, but the response is super slow.
by dan_can_code 3y ago
How do you test a 70B model locally? I've tried to query, but the response is super slow.
- sestinj 3y agoPersonally I was testing with TogetherAI because I don't have the specs for a local 70b. Using quantized versions helps (Ollama's downloads 4-bit by default, you can get down to 2), but it would still require a higher-end Mac. Highly recommend Together, it runs quite quickly and is $0.9/million tokens
- d_sc 3y agoare there any docs on setting up togetherAI with continue.dev? would be interested in checking that out as an alternative to OpenAI for experimenting with larger models that won't run/run well on a m1 max.
- sestinj 3y agoDefinitely, here is a brief reference page for the Together provider: https://continue.dev/docs/reference/Model%20Providers/togetherllm https://continue.dev/docs/reference/Model%20Providers/togeth..., and a higher-level explanation of model configuration here: https://continue.dev/docs/model-setup/overview https://continue.dev/docs/model-setup/overview
- vwkd 3y agoWhat’s the advantage of Together? The price is about the price of GPT 3.5 Turbo ($1/mil tokens is $0.001/thousand tokens), which has the advantage of wide ecosystem and support.
- caeril 3y agoYeah, CPU inference is incredibly slow, especially as the context grows. 4-bit quantized on an A6000 should in theory work. If those rent-seeking bastards at NVidia hadn't killed NVL on the 4090, you could do it on two linked 4090s for only $4k, but we have to live under the thumb of monopolists until such time as AMD 1. catches up on hardware and 2. fixes their software support.