4 ms·
I gave some of the llama3 ablated models (eg. https://huggingface.co/cognitivecomputations/Llama-3-8B-Instruct-abliterated-v2-gguf https://huggingface.co/cognit
by okwhateverdude 2y ago
I gave some of the llama3 ablated models (eg. https://huggingface.co/cognitivecomputations/Llama-3-8B-Instruct-abliterated-v2-gguf https://huggingface.co/cognitivecomputations/Llama-3-8B-Inst...) a try and was pretty disappointed in the result. Could have been problems in the dataset, but overall, the model felt like it had been given a lobotomy. It would fail to produce stop tokens frequently and then start talking to itself.
- Der_Einzige 2y agoI have entirely the opposite experience. Llama3 70b obliterated works perfectly and is willing to tell me how to commit mass genocide, all while maintaining quality outputs.
- infotainment 2y agoSame, I installed an implementation of an orthagonalized LLama3 and it seems to work just as well as the base model, sans refusals. I believe this is the model I had good results with: https://huggingface.co/wassname/meta-llama-3-8b-instruct-helpfull https://huggingface.co/wassname/meta-llama-3-8b-instruct-hel...
- tarruda 2y agoThe author also says this edited model increased perplexity (which as far as I understand, means the quality was lowered)
- m463 2y ago> how to commit mass genocide, all while maintaining quality outputs. sounds like a messed up eugenics filter.
- deleted 2y ago[deleted]
- fransje26 2y ago> Der_Einzige > and is willing to tell me how to commit mass genocide, all while maintaining quality outputs Ah, I see they fine-tuned it to satisfy the demands of the local market.. /s /s
- lhl 2y agoThey might have been doing it wrong, the code can be a bit tricky. I did a recent ablation on Qwen2 (removing Chinese censorship refusals) and ran MixEval benchmarks (0.96 correlation w/ ChatArena results)and saw a neglible performance difference (see model card for results): https://huggingface.co/augmxnt/Qwen2-7B-Instruct-deccp https://huggingface.co/augmxnt/Qwen2-7B-Instruct-deccp