4 ms·
Different quantizations can give you a big speedup if you've had "depressingly slow" issues. Even the slowest ones (that fit in RAM) will run at basically inter
by PostOnce 3y ago
Different quantizations can give you a big speedup if you've had "depressingly slow" issues. Even the slowest ones (that fit in RAM) will run at basically interactive speed, not instant, but also not "email speed". I have a laptop with a 2018 CPU and I'm working with them just fine.
Text generation style instead of chat style is another avenue that makes the feedback time not so annoying for a developer.
at 100ms/token, it's faster than most people type, I think. That's what you might get on an old laptop with a 7B model.
There's a useful leaderboard here to help you pick a model: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
It really depends on your task, lots and lots of natural language type tasks give great results, the models seem to have extensive knowledge of many fields. So for some kinds of Q&A bot (technical or not), for copy blurbs, for fiction, game NPCs, etc, the models (especially 13B and up) can be breathtaking, even moreso considering they run on bottom-dollar consumer hardware (I paid $250 for the laptop I'm developing on).
There are of course some things that neither the local LLMs nor GPT4 can do, like create useful OpenSCAD models :)
Things keep getting better, newer quantization methods give you more smarts in the same amount of RAM at basically the same speed -- the models are getting better, there are more permissively licensed ones now.
- RossBencina 3y agoHow are you running inference? GPU or CPU? I'm trying to use GPT4All (ggml-based) on 32 cores of E5-v3 hardware and even the 4GB models are depressingly slow as far as I'm concerned (i.e. slower than the GPT4 API, which is barely usable for interactive work). I'd be much obliged if you could point me at a specific quantized model on HF that you think is "fast" and I'll download it and try it out.
- wokwokwok 3y agoWhaaaaat, how are you getting 100ms per token on an 5 year old potato without a graphics card? Like, not vaguely hand wavey stuff, specifically, what model and what inference code? I get nothing like that performance for the 7B models, forget the larger models, using llama.cpp on a pc without an nvidia GPU.
- PostOnce 3y agoHere is a short test of a 7B 4bit model on an intel 8350U laptop with no AMD/Nvidia GPU. On that laptop CPU from 2017, using a copy of llama.cpp I compiled 2 days ago (just "make", no special options, no BLAS, etc): ./main -m models/WizardLM-7B-uncensored.ggmlv3.q4_0.bin -n 128 -s 99 -p "A short test for Hacker News:" llama_print_timings: sample time = 19.12 ms / 36 runs ( 0.53 ms per token, 1882.65 tokens per second) llama_print_timings: prompt eval time = 886.82 ms / 9 tokens ( 98.54 ms per token, 10.15 tokens per second) llama_print_timings: eval time = 5507.31 ms / 35 runs ( 157.35 ms per token, 6.36 tokens per second) and a second run: ./main -m models/WizardLM-7B-uncensored.ggmlv3.q4_0.bin -n 128 -s 99 -p "Sherlock Holmes favorite dinner was " llama_print_timings: sample time = 54.37 ms / 102 runs ( 0.53 ms per token, 1875.93 tokens per second) llama_print_timings: prompt eval time = 876.94 ms / 9 tokens ( 97.44 ms per token, 10.26 tokens per second) llama_print_timings: eval time = 16057.95 ms / 101 runs ( 158.99 ms per token, 6.29 tokens per second) at 158ms per token, if we guess a word is 2.5 tokens, then that's 151 words per minute, much faster than most people can type. On a $250 laptop. Isn't the future neat? the code I was running: https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp and the model: https://huggingface.co/TheBloke/WizardLM-7B-uncensored-GGML https://huggingface.co/TheBloke/WizardLM-7B-uncensored-GGML There are other models that may perform better, I'm going to be doing a lot of screwing around with OpenLLaMA this weekend.
- Rastonbury 3y agoThis is a exactly a case in point why people decide to pay OpenAI instead of rolling their own. I'm non-technical but have setup an image gen app based custom SD model using diffusers, so not entirely clueless. But for LLM I have no where idea where to start quickly. Finding a model on a leaderboard, download and setup then customising it and benchmarking is way too much time for me, I'll just pay for GPT4 if ever need to instead of chasing and troubleshooting to get some magical result. It'll be easier in the future I'm sure when an open model merges as the SD1.5 of LLM
- lhl 3y agoI've found https://gpt4all.io/ https://gpt4all.io/ to be the fastest way to get started. I've also started moving my notes to https://llm-tracker.info/ https://llm-tracker.info/ which should help make it easier for people getting started: https://llm-tracker.info/books/howto-guides/page/getting-started https://llm-tracker.info/books/howto-guides/page/getting-sta...