Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gpjt
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
Writing an LLM from scratch, part 32e – Interventions: the learning rate
(gilesthomas.com)
3 points
by
gpjt
7mo ago
|
0 comments
62.
▲
Writing an LLM from scratch, part 32d – Interventions: adding attention bias
(gilesthomas.com)
6 points
by
gpjt
8mo ago
|
0 comments
63.
▲
Writing an LLM from scratch, part 32c – Interventions: removing dropout
(gilesthomas.com)
1 points
by
gpjt
8mo ago
|
0 comments
64.
▲
Writing an LLM from scratch, part 32B – Interventions: gradient clipping
(gilesthomas.com)
2 points
by
gpjt
8mo ago
|
0 comments
65.
▲
Writing an LLM from scratch, part 32a – Interventions: training a baseline model
(gilesthomas.com)
1 points
by
gpjt
8mo ago
|
0 comments
66.
▲
Getting a Custom PyTorch LLM onto the Hugging Face Hub
(gilesthomas.com)
1 points
by
gpjt
8mo ago
|
0 comments
67.
▲
Writing an LLM from scratch, part 31 – the models are now on Hugging Face
(gilesthomas.com)
2 points
by
gpjt
9mo ago
|
0 comments
68.
▲
Writing an LLM from scratch, part 30 – digging into the LLM-as-a-judge results
(gilesthomas.com)
1 points
by
gpjt
9mo ago
|
0 comments
69.
▲
LLM from scratch, part 29 – using DDP to train a base model in the cloud
(gilesthomas.com)
2 points
by
gpjt
9mo ago
|
0 comments
70.
▲
by
gpjt
10mo ago
Not sure that every browser advertises English, but mine certainly does. However, as I'm in Portugal, many websites ignore what my browser says and send me to translated versions, I assume based on my IP. That causes problems because
71.
▲
by
gpjt
10mo ago
Awesome, thanks! I'm still doing trains on the big machines right now (hopefully will write up over xmas) but I think once I've worked out the sweet spot for memgatokens per dollar for this model, it's time to start tweaking
72.
▲
by
gpjt
10mo ago
I think the punctuation makes it clear -- imagine "How I invented Facebook. In 2001." The full stop in the middle of the sentence breaks it and makes you realise he's speaking figuratively.
73.
▲
by
gpjt
10mo ago
100%, I think there were weeks when I aged a year...
74.
▲
by
gpjt
10mo ago
Thanks re: gradient accumulation, I'm glad to hear my intuition was right! As part of the upcoming post I'm running the DDP train on A100s with 40 GiB and 80 GiB, H100s with 80 GiB, and B200s with 160 GiB, so I'll have at lea
75.
▲
by
gpjt
10mo ago
Hmm, interesting. With a batch size of 512 (8x B200s with 160 GiB each) I get worse results! Maybe there's a sweet spot somewhere in between.
76.
▲
by
gpjt
10mo ago
Exactly! If I can get it down to an hour or two (seems very plausible on an 8x H200 with 160 GiB VRAM per GPU, though those are almost never available on Lambda Labs), I'll do the experiments with dropout and the other possible causes
77.
▲
by
gpjt
10mo ago
OP here -- with a 112M model you should be able to get something worth playing with using 2.24B tokens. The Chinchilla heuristic is tokens = 20 x parameters. Obviously you cam get a better result by grinding through more tokens, but it will
78.
▲
by
gpjt
10mo ago
OK, early indicators support both you and Gemini quite strongly re: batch size. On my (somewhat ad-hoc) test dataset, I get losses like this: * OpenAI medium weights: 3.231 * OpenAI small weights: 3.500 * My locally trained model,
79.
▲
by
gpjt
10mo ago
OP here -- agreed! I tried to summarise (at least to my current level of knowledge) those 12-18 hours here: https://www.gilesthomas.com/2025/09/maths-for-llms
80.
▲
by
gpjt
10mo ago
OP here: one thing that surprised me in this experiment was that the model trained on the more curated FineWeb-Edu dataset was worse than the one trained on FineWeb. That is very counterintuitive to me.
81.
▲
by
gpjt
10mo ago
OP here -- thanks! I'm in the process of doing some trains using the same code plus DDP on big Lambda Labs machines, and (within the bounds of what I can afford) will hopefully have some interesting results about all of those shortly.
82.
▲
LLM from scratch, part 28 – training a base model from scratch on an RTX 3090
(gilesthomas.com)
540 points
by
gpjt
10mo ago
|
121 comments
83.
▲
by
gpjt
10mo ago
How much of that low survival rate is due to the condition they received the transplant, though? Conceivably a patient with "just" HIV might do better than one with eg. leukemia and HIV. That said, IIUC the whole stem cell transp
84.
▲
by
gpjt
11mo ago
Thanks for the reminder of a brilliant IT crowd moment!
85.
▲
Writing an LLM from scratch, part 27 – what's left, and what's next?
(gilesthomas.com)
1 points
by
gpjt
11mo ago
|
0 comments
86.
▲
Writing an LLM from scratch, part 26 – evaluating the fine-tuned model
(gilesthomas.com)
4 points
by
gpjt
11mo ago
|
0 comments
87.
▲
Writing an LLM from scratch, part 25 – instruction fine-tuning
(gilesthomas.com)
2 points
by
gpjt
1y ago
|
0 comments
88.
▲
Writing an LLM from scratch, part 24 – the transcript hack
(gilesthomas.com)
1 points
by
gpjt
1y ago
|
0 comments
89.
▲
Retro Language Models: Rebuilding Karpathy's RNN in PyTorch
(gilesthomas.com)
3 points
by
gpjt
1y ago
|
0 comments
90.
▲
Writing an LLM from scratch, part 23 – fine-tuning for classification
(gilesthomas.com)
1 points
by
gpjt
1y ago
|
0 comments
More ›