3 ms·
I also think that there is still room for improvement for embeddings based on other contexts as pointed in the blog entry. Another example from this year is lev
by cgravier 9y ago
I also think that there is still room for improvement for embeddings based on other contexts as pointed in the blog entry. Another example from this year is leveraging dictionary entries as external context -
http://aclweb.org/anthology/D17-1024 http://aclweb.org/anthology/D17-1024 ()
Selecting context words differently is also an option for improvement. Using dependency structures to "filter" out context window seems to work better than "filtering" using subsampling frequent words illustrate that there is room. We may see other solutions to select context words in the future, as a building block as it is. Especially lately with the StarSpace hype advocating the idea of general purpose - task-agnostic - embeddings.
Or we can also consider that the expected improvements are insignificant w.r.t. improvements with the model learnt on those embeddings for downstream tasks that may update embeddings especially for this task...
() disclaimer: I am a co author
- sebastianruder 9y agoThanks for the note, Christophe. I had missed your paper. I've added a short paragraph with regard to improving negative sampling by incorporating contextual information.
- cgravier 9y agoThank you Sebastian. Keep up the great work ! You will note that negative sampling improved by leveraging information on word pairs form dictionaries entries (we called it "controlled negative sampling") do help, though not much. It actually really depends in the rare words rate (see section 5.4, improvement ranges from 0.7% up to 10%). But I guess it is already an interesting, somehow counter-intuitive, observation. Another very interesting observation is that you can also choose to just clamp a general purpose dataset and expand it with external contextual information (meaning not using is for supervision but rather just collapse it at the end of the training corpus in a raw form [^]). In our case, we call those corpus : - corpus A : plain old wikipedia dump - corpus B : plain old wikipedia dump + dictionaries text collapse at the end of it. It sounds a bit naive : the latter part of the training corpus is really small w.r.t. the full wikipedia dump. Nonetheless, it has an significant impact on word similarity (see improvement in Table 2 to see how those training corpus influences representations learnt by word2vec, fasttext and dict2vec). (Related to the effect of the training corpus : https://arxiv.org/pdf/1507.05523v1.pdf https://arxiv.org/pdf/1507.05523v1.pdf) I mention this effect of training corpus content here since it sounds like an interesting info for the working natural language processing practitioners (get a mid size general training corpus, add as many contextual corpus as possible => may yield useful embeddings...). [^] to be entirely fair, this has been suggested to us by an anonymous reviewer, many thanks for him/her for pointing this out : I found the results surprising.
- cgravier 9y agoTypo from the blog : "to move related works" => "to move related words"