3 ms·
Indeed. I think "AI gold rush" sucks anyone with any skills in this area into it with relatively good pay, so there are no, or almost no people outside of big t
by danielEM 2y ago
Indeed. I think "AI gold rush" sucks anyone with any skills in this area into it with relatively good pay, so there are no, or almost no people outside of big tech and startups to counterbalance direction where it moves. And as a side note, big tech is and always was putting their agenda first in developing any tech or standards and that usually makes milking on investments as long as possible, not necessary moving things forward.
- llm_trw 2y agoThere's more to it than that. If you could train models faster, you’d be able to build larger, more powerful models that outperform the competition. The fact that Llama 3 is significantly over trained than what was considered ideal even three years ago shows there's a strong appetite for efficient training. The lack of progress isn’t due to a lack of effort. No one has managed to do this yet because no one has figured out how. I built 1-trit quantized models as a side project nearly a decade ago. Back then, no one cared because models weren’t yet using all available memory, and on devices where memory was fully utilized, compute power was the limiting factor. I spend much longer trying to figure out how to get 1-trit training to work and I never could. Of all the papers and people in the field I've talked to, no one else has either.
- sixfiveotwo 2y ago> I spend much longer trying to figure out how to get 1-trit training to work and I never could. What did you try? What were the research directions at the time?
- llm_trw 2y agoThis is a big question that needs a research paper worth of explanation. Feel free to email me if you care enough to have a more in-depth discussion.
- sixfiveotwo 2y agoSorry, I understand it was a bit intrusively direct. To bring some context, I toyed a little with neural networks a few years ago and wondered myself about this topic of training a so called quantized network (I wanted to write a small multilayer perceptron based library parameterized by the coefficient type - floating point or integer of different precision), but didn't implement it. Since you mentioned your own work in that area, it picked my interest, but I don't want to waste your time unnecessarily.
- llm_trw 2y agoSomeone posted a paper that I didn't know about, but goes through pretty much all the work I did in the space: https://news.ycombinator.com/item?id=42095999 https://news.ycombinator.com/item?id=42095999 It's missing the colourful commentary that I'd usually give, but alas, we can't have it all.
- sixfiveotwo 2y agothank you, that looks awesome.
- p1esk 2y agoPeople did care back then. This paper had jumpstarted the whole model compression field (which used to be a hot area of research in early 90s): https://arxiv.org/abs/1511.00363 https://arxiv.org/abs/1511.00363 Before that, in 2012, Alexnet had to be partially split into two submodels, running on two GPUs (using a form of interlayer grouped convolutions) because it could not fit in 3GB of a single 580 card. Ternary networks appeared in 2016. Unless you mean you actually tried to train in ternary precision - clearly not possible with any gradient based optimization methods.