3 ms·
Can someone please explain how this works to a software engineer who used to work with heuristically observable functions and algorithms? I'm having a hard time
by mercat 3y ago
Can someone please explain how this works to a software engineer who used to work with heuristically observable functions and algorithms? I'm having a hard time comprehending how a mix of experts can work.
In SE, to me, it would look like (sorting example):
- Having 8 functions that do some stuff in parallel
- There's 1 function that picks the output of a function that (let's say) did the fastest sorting calculation and takes the result further
But how does that work in ML? How can you mix and match what seems like simple matrix transformations in a way that resembles if/else flowchart logic?
- namibj 3y agoThe feed forward layer is essentially a differentiable key-value store. Similar to the attention layer, actually. So it just uses an attention mechanism like pre-selector to attend to only some experts. During inference, this cutoff is made a hard cutoff.
- mercat 3y agoThis is a very interesting approach. I know it may be too much to ask, but would you suggest any actual practical and hands-on workshops, playgrounds, or courses where I could practice using NN layers for stuff like that? For example, conditional/weighted selection of previous inputs, etc. It feels like I'm looking at ML programming from another angle.