Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
brrrrrm
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
121.
▲
by
brrrrrm
2y ago
At small input size, yes the MLP dominates compute. At large input attention matters more
122.
▲
by
brrrrrm
2y ago
You can quantize KV caches
123.
▲
by
brrrrrm
2y ago
HN isnt really the best space for LLM news - r/LocalLlama and twitter are much better. I think HN has some cultural issues with “AI” news
124.
▲
by
brrrrrm
2y ago
Maybe depth rather than parameter count.
125.
▲
by
brrrrrm
2y ago
Yep
126.
▲
by
brrrrrm
2y ago
bf16
127.
▲
by
brrrrrm
2y ago
I've been trying to think about how you'd amp up the batch size here. it's a bit tricky since the memory access would be way higher, but I think you can actually still save on compute by chunking things up in a clever way to
128.
▲
by
brrrrrm
2y ago
not sure why I got downvoted for asking this. seems like a reasonable thing in the ML space anyway here's a numerical simulation written in PyTorch for those who want to consider this algo on their projects before fully integrating: h
129.
▲
by
brrrrrm
2y ago
do you have a simple python impl? :)
130.
▲
by
brrrrrm
2y ago
companies that run these things care - they run at huge batch size and are compute bound in the limit
131.
▲
by
brrrrrm
2y ago
Yea. It’s one less hop through slow memory
132.
▲
by
brrrrrm
3y ago
It doesn’t. It simply trades compute efficiency by transposing matrix multiplications into “the future.” It doesn’t actually save FLOPs (uses more) and doesn’t work at large batch size
133.
▲
by
brrrrrm
3y ago
to add to this, "bagging" tracks are determined by how many shortcuts they have (ones that require good items to take, such as mushrooms or stars).
134.
▲
by
brrrrrm
3y ago
this one groups the characters/cars by identical stats: https://www.bettermk8dxbuilder.com
135.
▲
by
brrrrrm
3y ago
that one doesn't have mini-turbo, which has largely superseded acceleration as a stat (it correlates but isn't 1:1)
136.
▲
by
brrrrrm
3y ago
using AVX/FMA and unrolling loops does extremely little in the way of compiling to fast (>80% peak) GEMM code. These are very much intro steps that don't take into account many important ideas related to cache hierarchy, uop
137.
▲
by
brrrrrm
3y ago
Tree decoding exists and is being worked on by multiple groups. At the end of the day tree based methods are still just injecting some kind of interpretable prior.
138.
▲
by
brrrrrm
3y ago
Training is 3x the memory used by inference, and usually run at a much larger batch size
139.
▲
by
brrrrrm
3y ago
A provable way to recover convergence is to calculate the hessian. It’s computationally expensive but there are approximation methods.
140.
▲
by
brrrrrm
3y ago
Stuff like this always makes me think: what super structure do our brains form? Always end up concluding it’s the economy
141.
▲
by
brrrrrm
3y ago
The wings are doing a lot more work than the tail in that situation. For quick rotations in 3D, gyroscopic forces would probably be better to utilize than a tail
142.
▲
by
brrrrrm
3y ago
Cranking up the batch size kills convergence.
143.
▲
by
brrrrrm
3y ago
I predict knowing or caring about talking to a human via online anonymous chat is going to be a dated idea within five years. Might seem crazy today, but I think it’ll boil down to the conceptual equivalent of “is this text spell checked?”
144.
▲
by
brrrrrm
3y ago
Not really. It’s just a really simple hype-tech mashup and should be called out for it.
145.
▲
by
brrrrrm
3y ago
Are there any Reddit alternatives these days? Something like HN but more slightly more casual (image/video embeds, subreddits)
146.
▲
by
brrrrrm
3y ago
It looks like it's pretty resistant to quantization. ollama 4bit 7B doesn't work very well, but the 16bit 2B does
147.
▲
by
brrrrrm
3y ago
whoa, lots of negative reactions to a pretty mundane comment I made about the science of it all. It's overwhelmingly likely, from other studies, that brains differ, sure. But, the AI derived signal mentioned in this study might not
148.
▲
by
brrrrrm
3y ago
article about the paper: https://med.stanford.edu/news/all-news/2024/02/men-women-bra... "When the researchers tested the model on around 1,500 brain scans, it could almost always tell if the scan c
149.
▲
by
brrrrrm
3y ago
Recycled commentary on product changes made by an app trying to grow its ad revenue. Honestly this article could be AI generated
150.
▲
by
brrrrrm
3y ago
Not to slight the author, but Rust seems like quite a hard language to learn on the job.
More ›