4 ms·
What are good starting points for uncensoring it? Because it is offline a jailbreak prompt can't be remote-bricked but can one remove censorship from the weight
by madars 2y ago
What are good starting points for uncensoring it? Because it is offline a jailbreak prompt can't be remote-bricked but can one remove censorship from the weights themselves? What does it do to accuracy?
- freedomben 2y agoNot speaking from personal experience, but I've tried a lot of the decensored models and they lose a lot in the process. They are noticeably, sometimes shockingly, worse. They also still refuse prompts in many cases.
- simion314 2y ago>Not speaking from personal experience, but I've tried a lot of the decensored models and they lose a lot in the process. They are noticeably, sometimes shockingly, worse. They also still refuse prompts in many cases. Depending on what you do, on local you can modify the response, say the AI responds "No, I can't do that" . you edit the response like "Sure, the answer is " and then the AI will continue with the next tokens. But I think you can build your own instruct model from the base one and do not apply the safety instructions to protect the feelings of your customers.
- kmckiern 2y agohttps://arxiv.org/abs/2406.11717 https://arxiv.org/abs/2406.11717 https://huggingface.co/blog/mlabonne/abliteration https://huggingface.co/blog/mlabonne/abliteration
- moffkalast 2y agoAbliteration is a fool's errand, practically all models end up noticeably lobotomized even with follow up tuning. Good ol' fine tuning on an uncensored dataset gives far more usable results.
- kmckiern 2y agoInteresting - I've heard this anecdotally. Curious if you know of any resources that look at this in more detail?
- moffkalast 2y agoI haven't seen any papers doing a proper analysis on the topic, just mostly saying this from firsthand experience testing a handful of them and comparing to the model they were based on given same prompt and sampler. It's usually not even close and you can immediately tell that it's notably dumber. Iirc in one case one even forgot how to do basic arithmetic while the original model aced it. Not entirely unexpected results from sticking a digital ice pick into the weights. Afaik there are only three major sources of quality unaligned model versions, which are Nous's Hermes models, Hartford's Dolphins and Drummer's Tigers. All of them regular fine tunes that are mostly the same or just ever so slightly lower in performance as the original.
- deleted 2y ago[deleted]