4 ms·
Current LLMs cannot "actively participate in the real world" as humans do because they cannot actively learn from their interaction with the real world. Their w
by dimask 3y ago
Current LLMs cannot "actively participate in the real world" as humans do because they cannot actively learn from their interaction with the real world. Their weights are fixed. Their further training and finetuning is mediated by external mechanisms. It has nothing to do with what "people here" want.
- exoverito 3y agoSure, though it seems there are a number of near term paths for improvement. Short and long term memory mechanisms would go a long way towards active learning. Fine tuning could be iteratively performed through mechanisms like low rank adaptation. It's been hypothesized that this is one of the purposes of sleep, and humans are consolidating memories during dreams.
- candiodari 3y agoNot true. There's plenty of easy ways to update the weights if you want to do that. You could for example DPO based on another LLM's evalution of how a conversation is going, or you could use whether the conversation keeps going.
- staticman2 3y ago>>>"You could for example DPO based on another LLM's evalution of how a conversation is going," You can update the weights based on a single point of data (response was bad) but you probably can't usefully update a model that way.
- candiodari 3y agoI can't seem to find anyone actually usefully trying this. Do you know of any data?
- staticman2 3y agoNo but Google (for example) recommends fine-tuning based on at least 500 examples of question response pairs. DPO as far as I know requires a good and bad example. It's not a technique based on just saying "ai response is bad".