4 ms·
While that could be an explanation I am not fully sold on that idea. For example we don't get continuous streams of information on something like abstract math
by rdedev 2y ago
While that could be an explanation I am not fully sold on that idea. For example we don't get continuous streams of information on something like abstract math but we still dona decent job with it without going through millions of math problems.
If multi modality gets us through this phase then you are right in your analysis. Let's see what come out in the next 5 years
- mattnewton 2y agoI don’t mean to say that is _all_ that is separating us from the transformer architecture. I do think we’re more efficient than gradient descent, and sgd+attention is a function of how modern parallel processing works more than it is a function of any deep insight into the brain’s computation or learning methods. I just think we’re really making a silly comparison when we compare a stream of text to the multi sensory experiences humans have directly, reducing what we are doing to just reading words off a screen or page.