Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rain1
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
31.
▲
by
rain1
3y ago
unfortunately that chip is proprietary and undocumented, it's very difficult for open source programs to make use of. I think there is some reverse engineering work being done but it's not complete.
32.
▲
by
rain1
3y ago
wonderful! thank you
33.
▲
by
rain1
3y ago
> 1. Download the weights for the model you want to use, e.g. gpt4-x-vicuna-13B.ggml.q5_1.bin I think you need to quantize the model yourself from the float/huggingface versions. My understanding is that the quantization formats hav
34.
▲
by
rain1
3y ago
That is a crazy speedup!!
35.
▲
by
rain1
3y ago
I don't think the integrated GPU on that supports CUDA. So you will need to use CPU mode only.
36.
▲
by
rain1
3y ago
Tell us how it goes! Try different numbers of layers if needed. A good place to dig for prompt structures may be the 'text-generation-webui' commit log. For example https://github.com/oobabooga/text-generation
37.
▲
by
rain1
3y ago
You can offload only 10 layers or so if you want to run on a 4GB card
38.
▲
by
rain1
3y ago
my understanding is that the engine used (pytorch transformers library) is still faster than llama.cpp with 100% of layers running on the GPU.
39.
▲
by
rain1
3y ago
exciting! maybe we will see that land in llama.cpp eventually, who knows!
40.
▲
by
rain1
3y ago
I haven't tried that but https://github.com/abetlen/llama-cpp-python and https://github.com/r2d4/openlm exists
41.
▲
by
rain1
3y ago
4-bit quantization is to reduce the amount of VRAM required to run the model. You can run it 100% on CPU if you don't have CUDA. I'm not aware of any AMD equivalent yet.
42.
▲
by
rain1
3y ago
not at all, your question was really good so I added the answer to it to my gist to help everyone else. Sorry for the confusion I created by doing that!
43.
▲
by
rain1
3y ago
I'm sorry! I added this improvement based on that persons question!
44.
▲
by
rain1
3y ago
Apologies for that. I've added some extra micromamba setup commands that I should have included before! I've also added the git clone command, thank you for the feedback
45.
▲
by
rain1
3y ago
It's just for fun! These local models aren't as good as Bard or GPT-4.
46.
▲
by
rain1
3y ago
llama is a text prediction model similar to GPT-2, and the version of GPT-3 that has not been fine tuned yet. It is also possible to run fine tuned versions like vicuna with this. I think. Those versions are more focused on answering questi
47.
▲
Run Llama 13B with a 6GB graphics card
(gist.github.com)
618 points
by
rain1
3y ago
|
266 comments
48.
▲
by
rain1
3y ago
no
49.
▲
by
rain1
3y ago
this is not correct
50.
▲
by
rain1
3y ago
"I Crashed My Airplane" - TrevorJacob https://www.youtube.com/watch?v=vbYszLNZxhM
51.
▲
AI Scientists: Safe and Useful AI?
(yoshuabengio.org)
2 points
by
rain1
3y ago
|
0 comments
52.
▲
TEDx – Eliezer Yudkowsky – Unleashing the Power of Artificial Intelligence
(youtube.com)
2 points
by
rain1
3y ago
|
3 comments
53.
▲
Does prompt injection matter to AutoGPT?
(gist.github.com)
1 points
by
rain1
3y ago
|
0 comments
54.
▲
by
rain1
3y ago
python newbie here, why did messages change from [ to (
55.
▲
by
rain1
3y ago
There is no point in constructing a fixed template JSON object like that just to parse it again.
56.
▲
by
rain1
3y ago
thanks for clarifying that "pornhub" is an adult website
57.
▲
by
rain1
3y ago
> Together with Yann LeCun, and Yoshua Bengio, Hinton won the 2018 Turing Award for conceptual and engineering breakthroughs that have made deep neural networks a critical component of computing
58.
▲
by
rain1
3y ago
Is there a mod that adds PNP/NPN style transistors?
59.
▲
by
rain1
3y ago
why is it called serverless when it is a server
60.
▲
Pair Programming Experience with Bard
(gist.github.com)
2 points
by
rain1
3y ago
|
0 comments
More ›