4 ms·
>In reality, LLMs are already trained on the output of other LLMs That's a fact, but are there any studies on training a LLM on it's own output and not the out
by lewhoo 3y ago
>In reality, LLMs are already trained on the output of other LLMs
That's a fact, but are there any studies on training a LLM on it's own output and not the output of a different LLM ? For instance, chatgpt gets knowledge updates so I understand it must be retrained somehow. What happens if the retraining data contains large patches of its own output ? Has this scenario been explored ?
- _giorgio_ 3y agoDatasets are a dark subject because they contain illegal material. Most gpts, including chatGPT, are trained on the chatGPT outputs. Plus, a lot of chatGPT training material is hand picked by humans.