3 ms·
how hard is it to "update" an LLM like GPT-4 with recent data instead of being frozen in time at the training date? Obviously you could use the increased contex
by ren_engineer 4y ago
how hard is it to "update" an LLM like GPT-4 with recent data instead of being frozen in time at the training date? Obviously you could use the increased context size to work around this, but being able to augment the base model seems like it would be the ideal use case for many high value projects like a version of Copilot that knows a company's entire code base and knows about newer libraries and updates
- swyx 4y agoGPT has finetuning apis which you can use to update, but i think the massive context size is meant to show us you wont really need it most of the time. 25k words is really a lot of context. in the demo @gdb just dumped in the entire discord docs without breaking a sweat