3 ms·
The original article mentioned LLMs needing powerful abstractions this is basically the case with transformer networks, which is apparent when learning from sc
by lIIllIIllIIllII 3y ago
The original article mentioned LLMs needing powerful abstractions
this is basically the case with transformer networks, which is apparent when learning from scratch. The model seems to be going basically nowhere and totally useless until suddenly, at some random point after a bunch of learning cycles the weights find some minimum on the error surface and bam, suddenly the model can do things properly. And it's because the transformer has learned an abstraction that works for all of the input data in an attentional sense (think how you scan a sentence when reading). Not the best explanation but its from memory from a post I saw on HN a while back