48 ms·
Great post, thanks! Is there a reason why the training is started off with two separate matrices - the embedding and the context matrix? If the context matrix
by elexhobby 8y ago
Great post, thanks!
Is there a reason why the training is started off with two separate matrices - the embedding and the context matrix? If the context matrix is anyway discarded at the end, why not start and work with only the embedding matrix?