5 ms·
This is amazing work, but to me it highlights some of the biggest problems in the current AI zeitgeist, we are not really trying to work on any neuron or rulese
by retrofrost 3y ago
This is amazing work, but to me it highlights some of the biggest problems in the current AI zeitgeist, we are not really trying to work on any neuron or ruleset that isnt much different from the perceptron thats just a sumnation function. Is it really that suprising that we just see this same structure repeated in the models. Just because feedforward topologies with single neuron steps are the easiest to train and run on graphics cards does that really make them the actual best at accomplishing tasks? We have all sorts of unique training methods and encoding schemes that don't ever get used because the big libraries don't support them. Until, we start seeing real varation in the fundamental rulesets of neuralnets we are always just going to be fighting against the fact these are just perceptrons with extra steps.
- visarga 3y ago> Just because feedforward topologies with single neuron steps are the easiest to train and run on graphics cards does that really make them the actual best at accomplishing tasks? You are ignoring a mountain of papers trying all conceivable approaches to create models. It is evolution by selection, in the end transformers won.
- dartos 3y agoI mean RWKV seems promising and isn’t a transformer model. Transformers have first mover advantage. They were the first models that scaled to large parameter counts. That doesn’t mean they’re the best or that they’ve won, just that they were the first to get big (literally and metaphorically)
- tkellogg 3y agoYeah, I'd argue that transformers created such capital saturation that there's a ton of opportunity for alternative approaches to emerge.
- dartos 3y agoSpeak of the devil. Jamba just hit the front page.
- refulgentis 3y agoIt doesn't seem promising, a one man band has been doing a quixotic quest based on intuition and it's gotten ~nowhere, and it's not for lack of interest in alternatives. There's never been a better time to have a different approach - is your metric "times I've seen it on HN with a convincing argument for it being promising?" -- I'm not embarrassed to admit that is/was mine, but alternatively, you're aware of recent breakthroughs I haven't seen.
- dartos 2y agoRWKV has shown that you can scale RNNs to large parameter counts. The fact that one person (initially) was able to do it highlights how much low hanging fruit there is for non transformers. Also, the fact that a small number of people designed, trained, and published 5 versions of a perfectly serviceable (as in has decent summarizing ability. The biggest LLM use case) model which doesn’t have the time complexity of transformers is a big deal.
- retrofrost 3y agoJust because papers are getting published doesn't mean its actually gaining any traction. I mean we have known that time series of signals recieves plays a huge role in how bio neurons functionally operate and yet we have nearly no examples of spiking networks being pushed beyond basic academic exploration. We have known glial cells play a critical role in biological neural and yet you can probably count the number of papers that examine using an abstraction of that activity in neural net, on both your hands and toes. Neuroevolution using genetic algorithms has been basically looking for a big break since NEAT. Its the height of hubris to say that we have peaked with transformers when the entire field is based on not getting trapped in local maxima's. Sorry to be snippy, but there is so much uncovered ground its not even funny.
- gwervc 3y ago"We" are not forbidding you to open a computer, start experimenting and publishing some new method. If you're so convinced that "we" are stuck in a local maxima, you can do some of the work you are advocating instead of asking other to do it for you.
- Kerb_ 3y agoYou can think chemotherapy is a local maxima for cancer treatment and hope medical research seeks out other options without having the resources to do it yourself. Not all of us have access to the tools and resources to start experimenting as casually as we wish we could.
- erisinger 3y agoNot a single one of you bigbrains used the word "maxima" correctly and it's driving me crazy.
- vlovich123 3y agoAs I understand it a local maxima means you’re at a local peak but there may be higher maximums elsewhere. As I read it, transformers are a local maximum in the sense of outperforming all other ML techniques as the AI technique that gets the closest to human intelligence. Can you help my little brain understand the problem by elaborating? Also you may want to chill with the personal attacks.
- nicklecompte 3y agoHis point is that "evolution by selection" also includes that transformers are easy to implement with modern linear algebra libraries and cheap to scale on current silicon, both of which are engineering details with no direct relationship to their innate efficacy at learning (though indirectly it means you scale up the training data for more inefficient learning).
- wanderingbort 3y agoI think it is correct to include practical implementation costs in the selection. Theoretical efficacy doesn’t guarantee real world efficacy. I accept that this is self reinforcing but I favor real gains today over potentially larger gains in a potentially achievable future. I also think we are learning practical lessons on the periphery of any application of AI that will apply if a mold-breaking solution becomes compelling.
- foobiekr 3y ago"won" They barely work for a lot of cases (i.e., anything where accuracy matters, despite the bubble's wishful thinking). It's likely that something will sunset them in the next few years.
- victorbjorklund 3y agoThat is how evolution works. Something wins until something else comes along and win. And so on forever.
- Retric 3y agoEvolution generally favors multiple winners in different roles over a single dominate strategy. People tend to favor single winners.
- advael 3y agoI both think this is a really astute and important observation and also think it's an observation that's more true locally than of people broadly. Modern neoliberal business culture generally and the consolidated current incarnation of the tech industry in particular have strong "tunnel vision" and belief in chasing optimality compared to many other cultures, both extant and past
- imtringued 3y agoIn neoclassical economics, there are no local maxima, because it would make the math intractable and expose how much of a load of bullshit most of it is.
- foobiekr 3y agoYep. This. It’s impressive how communication is instantaneous, unimpeded, complete and transparent in economics. Those things aren’t even true in a 500 person company let alone an economy.
- 3y ago
- szundi 3y ago“end”
- antonvs 3y agoI’d say it’s more that transformers are in the lead at the moment, for general applications. There’s no rigorous reason afaik that it should stay that way.
- jjtheblunt 3y ago> in the end transformers won we're at the end?
- ldjkfkdsjnv 3y agoCannot understand people claiming we are in a local maxima, when we literally had an ai scientific breakthrough only in the last two years.
- xanderlewis 3y agoWhich breakthrough in the last two years are you referring to?
- ldjkfkdsjnv 3y agothe LLM scaling law
- 6gvONxR4sf7o 3y agoIf you had to reduce it to one thing, it's probably that language models are capable few shot and zero shot learners. In other words, training a model to simply predict the next word on naturally occurring text, you end up with an tool you can use for generic tasks, roughly speaking.
- xyzzy_plugh 3y agoIt turns out a lot of tasks are predictable. Go figure.
- ikkiew 3y ago> the perceptron thats just a sumnation[sic] function What would you suggest? My understanding of part of the whole NP-Complete thing is that any algorithm in the complexity class can be reduced to, among other things, a 'summation function'.
- blueboo 3y agoThe bitter lesson, my dude. http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html If you find a simpler, trainable structure you might be onto something Attempts to get fancy tried and died
- posix86 3y agoI don't understand enough about the subject to say, but to me it seemed like yes, other models have better metrics with equal model size i.t.o. number of neurons or asymptotic runtime, but the most important metric will always be accuracy/precision/etc for money spent... or in other words, if GPT requires 10x number of neurons to reach the same performance, but buying compute & memory for these neuros is cheaper, then GPT is a better means to an end.