3 ms·
A New Trick Could Block the Misuse of Open Source AI
- tithe 2y agoIf you have the ability to detect "harmful content" on the output side, shouldn't you be able to detect that material on the input side during training and exclude it so it never makes it into the model in the first place?
- Zambyte 2y agoThat is what Biderman suggests at the end of article. The problem with that is that models can produce novel information not found in the training data by correlating information that was not associated with each other directly in the training data. Realistically it should be up to the person running the inference to decide what information the model should be able to produce.
- throwaway888abc 2y agoTamper-Resistant Safeguards for Open-Weight LLMs https://arxiv.org/abs/2408.00761 https://arxiv.org/abs/2408.00761 p.s. if it's something tamper-resistant is still open ?
- infotainment 2y agoSad, but I have faith developers will figure out workarounds.