Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
gpjt
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
91.
▲
by
gpjt
1y ago
To be fair to the OpenAI team, if read in context the situation is at worst ambiguous. The deleted tweet that the article is about said "GPT-5 just found solutions to 10 (!) previously unsolved Erdös problems, and made progress on 11 o
92.
▲
Writing an LLM from scratch, part 22 – training our LLM
(gilesthomas.com)
254 points
by
gpjt
1y ago
|
10 comments
93.
▲
Revisiting Karpathy's 'Unreasonable Effectiveness of Recurrent Neural Networks'
(gilesthomas.com)
2 points
by
gpjt
1y ago
|
0 comments
94.
▲
Writing an LLM from scratch, part 21 – perplexed by perplexity
(gilesthomas.com)
1 points
by
gpjt
1y ago
|
0 comments
95.
▲
by
gpjt
1y ago
This is a great post on many levels, but what struck me as particularly clever was the use of lm_head to decode the outputs of earlier layers. That linear layer is only trained to decode the output of the last layer, so intuitively it might
96.
▲
Writing an LLM from scratch, part 20 – starting training, and cross entropy loss
(gilesthomas.com)
41 points
by
gpjt
1y ago
|
3 comments
97.
▲
How Do LLMs Work?
(gilesthomas.com)
2 points
by
gpjt
1y ago
|
1 comments
98.
▲
by
gpjt
1y ago
Post author here. I agree 100%! The post is the basic maths for people digging in to how LLMs work under the hood -- I wrote a separate one for non-techies who just want to know what they are, at https://www.gilesthomas.com/
99.
▲
by
gpjt
1y ago
Check the first link in the parent comment, it's a link to the book.
100.
▲
The maths you need to start understanding LLMs
(gilesthomas.com)
616 points
by
gpjt
1y ago
|
120 comments
101.
▲
What AI chatbots are doing under the hood
(gilesthomas.com)
2 points
by
gpjt
1y ago
|
0 comments
102.
▲
LLM from scratch, part 18 – residuals, shortcut connections, and the Talmud
(gilesthomas.com)
2 points
by
gpjt
1y ago
|
0 comments
103.
▲
The fixed length bottleneck and the feed forward network
(gilesthomas.com)
1 points
by
gpjt
1y ago
|
0 comments
104.
▲
Writing an LLM from scratch, part 17 – the feed-forward network
(gilesthomas.com)
8 points
by
gpjt
1y ago
|
0 comments
105.
▲
Writing an LLM from scratch, part 16 – layer normalisation
(gilesthomas.com)
1 points
by
gpjt
1y ago
|
0 comments
106.
▲
by
gpjt
1y ago
Congrats! Amazing feeling, isn't it :-)
107.
▲
Leaving PythonAnywhere
(gilesthomas.com)
3 points
by
gpjt
1y ago
|
0 comments
108.
▲
Writing an LLM from scratch, part 15 – from context vectors to logits
(gilesthomas.com)
7 points
by
gpjt
1y ago
|
0 comments
109.
▲
by
gpjt
1y ago
This, 100%. A full-stack engineer will likely have at least a solid understanding of the HTTP protocol, HTTPS, WebSockets, the interface layer between the frontend server and their chosen Web webdev stack, and so on. Then a more general und
110.
▲
Writing an LLM from scratch, part 14 – the complexity of self-attention at scale
(gilesthomas.com)
1 points
by
gpjt
1y ago
|
0 comments
111.
▲
by
gpjt
1y ago
As the author of the original post above, let me say that if that's word salad, it's a Michelin star salad. Just the right mix of lettuce and tomato, and the dressing is spot on :-) Seriously, though, differentiable hash tables is
112.
▲
by
gpjt
1y ago
Author of the post here -- I'm being careful not to do that. My posts are more about filling in the gaps; they're covering the things that aren't mentioned. The book's target audience is, I think, people with a bit more
113.
▲
by
gpjt
1y ago
Author here: I endorse this comment ;-) That's definitely the route I've optimised for for reading the series.
114.
▲
Writing an LLM from scratch, part 13 – attention heads are dumb
(gilesthomas.com)
351 points
by
gpjt
1y ago
|
67 comments
115.
▲
by
gpjt
1y ago
Another one leaving for Porkbun here.
116.
▲
by
gpjt
1y ago
Huh, I was thinking the same thing, and was wondering whether it was just moving to London. Could be both, I suppose.
117.
▲
by
gpjt
1y ago
As co-founder of Resolver Systems -- we tried but ultimately failed to take on Excel with a Python-enabled equivalent back in 2007 -- and current employee at Anaconda (providing Python in Excel) I really do hope you get this one to work. Ex
118.
▲
Writing an LLM from scratch, part 12 – multi-head attention
(gilesthomas.com)
3 points
by
gpjt
1y ago
|
0 comments
119.
▲
Writing an LLM from scratch, part 11 – batches
(gilesthomas.com)
2 points
by
gpjt
1y ago
|
0 comments
120.
▲
by
gpjt
1y ago
Interesting. As you say, that certainly makes sense for mammala. But I'd be interested in knowing what mechanisms you might conjecture for birds, where pretty much all foetal development happens inside the egg, separated from the mot
More ›