4 ms·
Depends on the definition of "knowledge"; there's a lot of factors that go into it. Some of the common approaches are continued/continual pretraining and model
by ijk 1y ago
Depends on the definition of "knowledge"; there's a lot of factors that go into it. Some of the common approaches are continued/continual pretraining and model editing (https://arxiv.org/pdf/2502.12598 https://arxiv.org/pdf/2502.12598).
* Models are bad at learning that A=B implies B=A, let alone more complicated relations; augmenting the dataset with multiple examples with different phrasing/perspectives is important (https://arxiv.org/abs/2404.00213 https://arxiv.org/abs/2404.00213). The frequency that a relation occurs in the dataset affects the results (https://arxiv.org/html/2504.09597v2 https://arxiv.org/html/2504.09597v2).
* You have to be able to balance preserving existing knowledge against the new knowledge (https://arxiv.org/abs/2502.14502 https://arxiv.org/abs/2502.14502). There are techniques like making sure your data mix corresponds to the original training data, but new data is primed by existing data so it gets complicated (https://arxiv.org/abs/2504.09522 https://arxiv.org/abs/2504.09522).
* Curriculum training (a la Phi) can be quite effective for training knowledge into base models at the very least.
* Continued pretraining is much more difficult than most finetuning, though it is possible (https://unsloth.ai/blog/contpretraining https://unsloth.ai/blog/contpretraining).
* Model editing of individual facts is possible but tricky because everything is interconnected but the model isn't great at figuring out reciprocal relationships (https://arxiv.org/abs/2310.16218 https://arxiv.org/abs/2310.16218). There's been some slow progress, though I find that few people are aware that it is even possible, despite the progress that has been made (https://github.com/zjunlp/KnowledgeEditingPapers https://github.com/zjunlp/KnowledgeEditingPapers).
The keywords you want are knowledge injection, domain adaptation, continual pretraining, model editing.
- simianwords 1y agoThis is exactly what I was talking about. I wonder why no one has tried to inject a critical code repository (at least 1 million LOC) and compare to common RAG methods? The ones you have shown here are nice and simple like world cup statistics. Maybe we are nowhere near solving such complicated scenarios?