Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
chessgecko
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
chessgecko
2y ago
This is the sparsest model thats been put out in a while (maybe ever, kinda forget the shapes of googles old sparse models). This probably wont be a great tradeoff for chat servers, but could be good for local stuff if you have 512GB of ram
32.
▲
by
chessgecko
2y ago
I think its almost certainly using at least two experts per token. It helps a lot during training to have two experts to contrast when putting losses on the expert router.
33.
▲
by
chessgecko
2y ago
For this comparison the generation of chip doesn’t really matter because the llm decode (which is the costly step) barely uses any of the perf and just needs the model weights to fit in memory
34.
▲
by
chessgecko
2y ago
My read of patent one is that they basically created DNS for storage. But DNS was invented in 1983 so I'm not really sure what was novel here other than pointing it at data, which uses a few extra headers, ala my comment. Even if there
35.
▲
by
chessgecko
2y ago
Links to two of the patents that were infringed https://patents.google.com/patent/US7103640B1/en https://patents.google.com/patent/US7233978B2/en I really hope Kove loses, I don't k
36.
▲
by
chessgecko
2y ago
they didn't use GDDR cause they wanted the memory capacity which is really important for recommendation models. But I totally agree that this is a sort of perfect cost/perf per watt point for a home setup. I really hope they do it
37.
▲
by
chessgecko
2y ago
its gotta be that 2/4 sparsity that everyone has, but I haven't seen used anywhere right? If they put it in though they must be using it, but I'm not sure for what. And without details I think its a good bet that int8 is the
38.
▲
by
chessgecko
2y ago
Also its at 90 watts vs 900 watts for gaudi 3, the flops/mem bw per watt is much more comparable.
39.
▲
by
chessgecko
2y ago
I thought MTIA v2 would use the mx formats https://arxiv.org/pdf/2302.08007.pdf , guess they were too far along in the process to get it in this time. Still this looks like it would make for an amazing prosumer home ai
40.
▲
by
chessgecko
3y ago
They could, I know they wont, but they wouldn't lose money on the parts
41.
▲
by
chessgecko
3y ago
I feel a little misled by the speedup numbers. They are comparing lower batch size h100/200 numbers to higher batch size gaudi 3 numbers for throughput (which is heavily improved by increasing batch size). I feel like there are some in
42.
▲
by
chessgecko
3y ago
I think you're right on the price, but just to give some false hope. I think newish hbm (and this is hbm2e which is a little older) is around $15/gb so for 128 gb thats $1920. There are some other cogs, but in theory they could se
43.
▲
by
chessgecko
3y ago
You could, but the memory bandwidth wouldn’t be amazing unless you had a lot of sticks and it would end up getting pretty expensive
44.
▲
by
chessgecko
3y ago
Going above 24GB is probably not going to be cheap until gddr7 is out, and even that will only push it to 36gb. The fancier stacked gddr6 stuff is probably pretty expensive and you can’t just add more dies because of signal integrity issues
45.
▲
by
chessgecko
3y ago
There’s some fancier stuff too like techniques that take into account where recent tokens were drawn from in the distribution and update either the top_p or the temperature so that sequences of tokens have a minimum unlikeliness. Beam searc
46.
▲
by
chessgecko
3y ago
Probably the same analytics/funnel bs that's everywhere these days. Though in this case the users are probably vcs not buyers.
47.
▲
by
chessgecko
3y ago
It's a shame they won't make a pcie mi300x. It has about the same amount of compute/memory as this and if the rumor mill is right it would cost almost the same.
48.
▲
by
chessgecko
3y ago
Tensorflow and keras have gotten better, but pytorch historically had better flexibility than keras and was much easier to debug/develop in than tensorflow.
49.
▲
by
chessgecko
3y ago
also just to add, I think the 1.58 bit is mostly faster for inference because training still had to multiply a lot of floating point gradients by integer activations, hold floating point weights/gradients for round, and deal with norms
50.
▲
by
chessgecko
3y ago
The problem is that it’s probably often not a lot cheaper. Most of the high end gpus have comparatively little bandwidth over pcie (that you’d need to use to store the context on a nvme for example). The cost there would scale with length t
51.
▲
by
chessgecko
3y ago
This is wrong, being memory bound or not has to do with the dimensions of the matrices being multiplied (if you’re on tensor cores). https://docs.nvidia.com/deeplearning/performance/dl-performa... Some of the thin
52.
▲
by
chessgecko
3y ago
It takes a 2/1.5bit model, groups parameters together then exploits a lack of entropy in the parameters to compress it a bit like text compression. It was only below 1bit for the ultra large model, guess the smaller ones weren’t quite
53.
▲
by
chessgecko
3y ago
Not sure if I read it correctly, but it seems like the skip connections are kinda still present in the skipless block because they added I after the softmax and v= the previous hidden state. Still a cool paper if they really managed to get
54.
▲
by
chessgecko
3y ago
People who dislike things are just so much more vocal than people who like them. I've used Huggingface extensively, they are trying to do a lot, but its always been the most convenient/flexible for my finetuning use cases. Thank y
55.
▲
by
chessgecko
3y ago
He refers all over the blog post to an "error" in attention. specifically says The problem with using softmax is that it forces each attention head to make an annotation, even if it has no information to add to the output vector.
56.
▲
by
chessgecko
3y ago
I guess yeah I was mostly responding to Now it’s possible that softmax should be replaced wholesale, but it’s worked pretty well for the most part, except for this one wee little bug that prevents attention heads from saying nothing. So I
57.
▲
by
chessgecko
3y ago
It wasn't really the goal of my experiment to fix this issue for sure, I was trying to see if you could improve attention by decoupling the key used by a position for itself and for future tokens. Open to being wrong here, but wouldn&#
58.
▲
by
chessgecko
3y ago
I was just looking at doing this in pretraining, so I was looking at pretraining losses. The difference was within the range of usual noise so I didn't keep trying.
59.
▲
by
chessgecko
3y ago
I think maybe its because he didn't have experimental results that show that it worked. Not a knock against the author, there are just so many things that seem like good ideas that don't end up working well in practice, a paper li
60.
▲
by
chessgecko
3y ago
I ran an experiment like this and in my setting it didn't help. Not saying there may not have been a bug or something, but I think attending over the current position sort of solves this problem. IE when it should not speak it just emi
More ›