Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zhisbug
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
Important and *MUST-KNOW* techniques for a 2023 LLM serving system
(twitter.com)
1 points
by
zhisbug
3y ago
|
0 comments
32.
▲
by
zhisbug
3y ago
oops typo
33.
▲
by
zhisbug
3y ago
yep, it is agonistic to 4-bit. You can deploy a 4-bit model and still use vllm + pagedattention to double or even triple your serving throughput.
34.
▲
by
zhisbug
3y ago
This really depends on what GPUs you use. If you GPUs has very small amount of memory, vLLM will help more. vLLM addresses the memory bottleneck for saving KV caches and hence increases the throughput.
35.
▲
by
zhisbug
3y ago
could you provide any reference? Is it a variant of ELO?
36.
▲
Fastchat-T5: 4x smaller but more powerful than Dolly-v2, commercial use ready
(twitter.com)
7 points
by
zhisbug
3y ago
|
1 comments
37.
▲
by
zhisbug
3y ago
Released by Lmsys.org who made Vicuna
38.
▲
by
zhisbug
4y ago
And I think the problem of taking the roles of users in vicuna is caused by this bug: https://github.com/lm-sys/FastChat/commit/1bb234265d16bdfd50... which has been fixed recently. Lmsys are launching new tra
39.
▲
by
zhisbug
4y ago
My point is that I am not aware of any official 4-bit quantization version (delta or weights) by lmsys so it might too early to draw your conclusion that vicuna finetuned llamas degenerates a lot of performance at 4 bit but others are fine.
40.
▲
by
zhisbug
4y ago
it is really a matter of having faith on pytorch (or JAX) or on third-party cross-platform supports like llama-cpp. Apparently pytorch reduces a lot of complexity and grows extremely faster on cross-platform supports. And, PyTorch does so w
41.
▲
by
zhisbug
4y ago
no, the requirement on a particular HF commit has been fixed. It is no longer needed.
42.
▲
by
zhisbug
4y ago
Lmsys hasn't released any official 4-bit version. It might be a better idea to wait for the official 4-bit version. But it is interesting to learn that the third-party 4bit version has performance degeneration.
43.
▲
by
zhisbug
4y ago
No llama.cpp nor any compilation complexity. Run with two Python commands!
44.
▲
by
zhisbug
4y ago
https://github.com/facebookresearch/llama/pull/184
45.
▲
by
zhisbug
4y ago
lol weights are all you need
46.
▲
by
zhisbug
4y ago
GPT4 + retrieval might be the fastest path. But quality not guaranteed, and assuming you do not mind uploading all your private info to openAI. This project might be the best option where you can finetune an LLM on your data and keep the mo
47.
▲
by
zhisbug
4y ago
Yeah that might work but this model wasn’t tuned with lora
48.
▲
by
zhisbug
4y ago
The recipe indeed may not be a bad idea lol
49.
▲
by
zhisbug
4y ago
and I do think this is a good effort
50.
▲
by
zhisbug
4y ago
Exactly. but if you want some commercial use, still lacking good quality models.
51.
▲
by
zhisbug
4y ago
but it is indeed difficult to eval chatbots and LLM esp. considering most of them have actually seen the Internet data at least once.
52.
▲
by
zhisbug
4y ago
but how could you guarantee the pretrained model haven't seen those benchmarks? And the baselines you are comparing to (chatgpt, bard) haven't as well? Cuz those benchmarking datasets are also collected from Internet right?
53.
▲
by
zhisbug
4y ago
> We plan to release the model weights by providing a version of delta weights that build on the original LLaMA weights, but we are still figuring out a proper way to do so.
54.
▲
by
zhisbug
4y ago
yeah this model is primarily tuned based on English data
55.
▲
by
zhisbug
4y ago
then do you have better way to more rigorously evaluate chatbot at the presence of LLMs like ChatGPT trained on almost all Internet data?
56.
▲
by
zhisbug
4y ago
very hopeful!
57.
▲
by
zhisbug
4y ago
why do you think so?could you elaborate?
58.
▲
by
zhisbug
4y ago
or Andes which means you breeds all llamas lol
59.
▲
by
zhisbug
4y ago
maybe you can restart the series with camels instead of llamas
60.
▲
by
zhisbug
4y ago
Serve OPT-175B with Alpa using commodity GPUs: https://alpa-projects.github.io/tutorials/opt_serving.html If you have difficulties accessing 8x 80GB A100 (or AWS P4) this might be a good use case.
More ›