Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
MakazhanAlpamys
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
Show HN: A Soup-inspired runtime for a real fruit-fly connectome =)
(github.com)
4 points
by
MakazhanAlpamys
23d ago
|
0 comments
2.
▲
by
MakazhanAlpamys
24d ago
yoo, it's good? how is it work?
3.
▲
by
MakazhanAlpamys
26d ago
good
4.
▲
by
MakazhanAlpamys
2mo ago
i don't follow gpu prices. the cheapest one is the one you already have. more vram for the money is what i'd look for. usually used :)
5.
▲
by
MakazhanAlpamys
2mo ago
you're right. it was the second, not the first. i already admitted that earlier in the thread :)
6.
▲
by
MakazhanAlpamys
2mo ago
Those are format examples and test fixtures. Five to ten rows each. Not training data. You did spot a real problem though. Eight configs in `examples/configs` pointed at those fixtures as training data. Seven were still on the old sche
7.
▲
by
MakazhanAlpamys
2mo ago
It isn't. Kazakh and Russian. I said this further down but that comment is dead so you would not have seen it. The later replies are mine, written by me.
8.
▲
by
MakazhanAlpamys
2mo ago
Do not buy a 4 GB card for this. Mine is an RTX 3050 Laptop, I picked it because it is boring hardware that many people already have. If you are buying, buy VRAM. At 0.5B where I could measure both, resident training was 1.43x faster than s
9.
▲
by
MakazhanAlpamys
2mo ago
Because streaming only removes the decoder stack. The embeddings and lm_head stay resident, that is 2.10 GB of the 3.32 GB peak on 8B. And the logits tensor scales with batch x seq x vocab, not with depth. So it goes from "whole model
10.
▲
by
MakazhanAlpamys
2mo ago
Depends what you change. Format or style, few hundred rows is often enough. A task the model already half knows, few thousand. New facts is where people waste a week. The model comes back wrong in a new way. Use RAG for facts.
11.
▲
by
MakazhanAlpamys
2mo ago
Those are mine. I used an LLM for my replies and that was a bad call, I said so further down. Writing them myself now.
12.
▲
by
MakazhanAlpamys
2mo ago
Author here. The constraint everyone works around is that the frozen base has to fit in VRAM. But during LoRA the base is frozen — read, never written. It doesn't need to live in VRAM, it needs to arrive before the matmul that uses it.
13.
▲
Show HN: Fine-tune an 8B model on a 4 GB laptop GPU
(github.com)
139 points
by
MakazhanAlpamys
2mo ago
|
28 comments