3 ms·
https://arxiv.org/abs/2406.11717 https://arxiv.org/abs/2406.11717 https://huggingface.co/blog/mlabonne/abliteration https://huggingface.co/blog/mlabonne/abliter
by kmckiern 2y ago
https://arxiv.org/abs/2406.11717 https://arxiv.org/abs/2406.11717
https://huggingface.co/blog/mlabonne/abliteration https://huggingface.co/blog/mlabonne/abliteration
- moffkalast 2y agoAbliteration is a fool's errand, practically all models end up noticeably lobotomized even with follow up tuning. Good ol' fine tuning on an uncensored dataset gives far more usable results.
- kmckiern 2y agoInteresting - I've heard this anecdotally. Curious if you know of any resources that look at this in more detail?
- moffkalast 2y agoI haven't seen any papers doing a proper analysis on the topic, just mostly saying this from firsthand experience testing a handful of them and comparing to the model they were based on given same prompt and sampler. It's usually not even close and you can immediately tell that it's notably dumber. Iirc in one case one even forgot how to do basic arithmetic while the original model aced it. Not entirely unexpected results from sticking a digital ice pick into the weights. Afaik there are only three major sources of quality unaligned model versions, which are Nous's Hermes models, Hartford's Dolphins and Drummer's Tigers. All of them regular fine tunes that are mostly the same or just ever so slightly lower in performance as the original.