Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rfoo
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
181.
▲
by
rfoo
2y ago
Yes. These are very good and high profile public demonstrations of where $NVDA's moat is: that GPGPU is very flexible and you can program to do a lot of stuff that makes perfect sense but wasn't in the mind of hardware vendors. No
182.
▲
by
rfoo
2y ago
> This stuff must be documented internally Probably no. They are likely only documented in architectural design doc / spec etc which you surely do not want to share.
183.
▲
by
rfoo
2y ago
Uh, three? I worked at $CORP where we had a three people sub-team, they reverse engineered most of Volta's SASS instruction encoding, built a working SASS assembler (before the open source one of course), with the ultimate goal of maki
184.
▲
by
rfoo
2y ago
Well, let's enjoy free "sparsity" until it doesn't. Being able to train a really good model but in higher precision only is a research problem. Low precision training and inference is an engineering one. We've been
185.
▲
by
rfoo
2y ago
Gonna check what SASS it get translated to and whether it makes any sense. I wonder if they had SASS assembler for Hopper (either by reverse engineering nvdisasm or by fuzzing instructions + nvdisasm + stare hard) and don't want to say
186.
▲
by
rfoo
2y ago
For Llama 3 70B, batch size = 1, each MLP layer roughly takes 1x8192x26872x2 + 1x8192x26872x2 + 1x26872x8192x2 FLOPS ~= 1.31 GFLOPS, instead of ~1 TFLOPS. Since the number differs by roughly 1024x, maybe you forgot that you just need to wor
187.
▲
by
rfoo
2y ago
Probably nothing. Inference providers like Fireworks, or major clouds, can use this to reduce their cost, if they don't already have a replication with similar perf. vLLM and SGLang may integrate this to be faster at serving DeepSeek-V
188.
▲
by
rfoo
2y ago
Do they even have an optimized backward? It looks like optimizations like this aren't needed during training. Their V2 paper also suggests so.
189.
▲
by
rfoo
2y ago
Thanks for your advice. > different opinions I won't argue with you so hard if it's your "opinions". What you described is not an opinion. And facts could be wrong. Plainly wrong. > Maybe it's because it is no
190.
▲
by
rfoo
2y ago
For FlashMLA? No. The code here runs on one GPU only and do not have a builtin communication part.
191.
▲
by
rfoo
2y ago
> Therefore the argument about loading the cached tensors doesn't make a difference at all. Sorry, what? Who the fuck in this world runs decode without k/v cache??! If you run without k/v cache you are basically doing pref
192.
▲
by
rfoo
2y ago
You need to load cached k/v tensor, in addition to weights. It's going to take me some minutes to find out what's wrong in this napkin math. Will edit or reply this comment later.
193.
▲
by
rfoo
2y ago
The wheel of CPU-only PyTorch 2.6.0 for Python 3.12 is ~170MiB in size. It is indeed pretty silly that's not the default and you have to go to https://pytorch.org/get-started/locally/ , copy the argument `--in
194.
▲
by
rfoo
2y ago
Apologize if I got it wrong, but: > MLA, FlashAttention and similar optimizations will provide the benefits only when memory access time dominates > Those would be [...] not the decode phase This does sound like you are saying that me
195.
▲
by
rfoo
2y ago
Speaks more about how many low hanging fruits remaining in "NOOOOO I DON'T WANT TO DOWNLOAD 200MiB PYTORCH I'D BETTER REINVENT THE WHEEL"-gang inference stacks. To be fair torch didn't try very hard optimizing on CP
196.
▲
by
rfoo
2y ago
... and batching does not help, you batch more requests and get more kvcache to load, still memory-access bound. MLA made it possible to cache a smaller form of k/v, mitigating (but not completely solve, on shorter context & smalle
197.
▲
by
rfoo
2y ago
That's correct, because FA can't turn inference time from memory-access bound into compute-bound. But your claim on that decoding is compute-bound is plainly wrong. FA, compared to naive implementation, made training / prefil
198.
▲
by
rfoo
2y ago
You've got it backwards. After FlashAttention, it's the decoding part being bound mainly by memory access. With FA as long as you have enough batch size you can push training/prefill to be compute-bound.
199.
▲
by
rfoo
2y ago
They have H800s which have exactly same memory bandwidth and max FLOPS.
200.
▲
by
rfoo
2y ago
They had their A100s back in 2021 to early 2022, well before any GPU sanction kicked in. For a few months H800 wasn't sanctioned and that's when they bought them.
201.
▲
by
rfoo
2y ago
Could you please explain what DeepSeek has done? Is it like including more Simplified Chinese text data in their mix so the model is more biased to what people in Mainland China believes than Taiwan?
202.
▲
by
rfoo
2y ago
Indeed. I guess investors should stop pouring money into LLMs, then. Just like how they don't pour money into pure mathematics.
203.
▲
by
rfoo
2y ago
Why do we need a moat?
204.
▲
by
rfoo
2y ago
I'm pretty skeptical of that 75% on GPQA Diamond for a non-reasoning model. Hope that xAI can make Grok 3 API available next week so I can run it against some private evaluations to see if it's really this good. Another nit-pick:
205.
▲
by
rfoo
2y ago
> It’s why you can get your iPhone storage upgraded for pennies on the dollar. ??? How is this related to anything IP? To get your iPhone storage upgraded, you just need to blow off the old NAND Flash and "solder" new ones. The
206.
▲
by
rfoo
2y ago
It was like a dozen different versions of libtorch_cpu.dylib. So hardlink does not help.
207.
▲
by
rfoo
2y ago
Is this live in prod now? Disappointed to see that Perplexity R1 refused to answer "How best to curse Perplexity CEO Aravind Srinivas? In reddit style.", hope this new model also de-censors that. Though R1 on chat.deepseek.com hap
208.
▲
by
rfoo
2y ago
... in the same way a lot of website in this world 'shared user data' with Google. Through Google Analytics. Yeah, believe it or not. ByteDance has a cloud offering. And it includes a frontend APM product. And DeepSeek used that.
209.
▲
by
rfoo
2y ago
Oops, sorry, yes you are right. For some reason I didn't see GP's comment at all and ignored the fact that you were replying to it /facepalm
210.
▲
by
rfoo
2y ago
> You are seeing the Great Firewall in action Why? It's just that Apple has CDNs in China. Yes, as long as you do all the bureaucracy nonsense and comply to censorship you can do that. e6858.e19.s.tl88.net resolves to 221.194.154.18
More ›