3 ms·
There are versions of Qwen3.8-27B that are unrestricted and available from hugging face. "It will comply with harmful, unethical, offensive, or illegal request
by dantudor 1mo ago
There are versions of Qwen3.8-27B that are unrestricted and available from hugging face.
"It will comply with harmful, unethical, offensive, or illegal requests that the original Qwen3.8-27B would refuse. It has no meaningful built-in guardrails."
- radlad 1mo ago> What makes this build different is the word before FP8: uncensored. We applied abliteration — orthogonalizing the refusal direction out of the residual stream — to remove the model's safety-alignment refusals. The result is a model that will comply with requests the original would refuse. Surely this has unintended side effects on output quality?
- andsoitis 1mo ago> > What makes this build different is the word before FP8: uncensored. We applied abliteration — orthogonalizing the refusal direction out of the residual stream — to remove the model's safety-alignment refusals. The result is a model that will comply with requests the original would refuse. > Surely this has unintended side effects on output quality? Can you help me understand why that's the case?
- willy_k 1mo agoBecause deleting model weights after training is likely to cause knock-on effects in model knowledge and/or behavior. Targetting it might mitigate this but it’s a) not guaranteed that only censor-ey parameters get removed, and b) likely that removing those parameters still has effects on the effectiveness of related parameters.
- jszymborski 1mo agoThe weights aren't deleted, it's just additional fine tuning, is my understanding.
- naasking 1mo agoThere is no question model quality is degraded by this though.
- InvertedRhodium 1mo agoIt's altered, sure. I think inherent degradation is a step too far though.
- badsectoracula 1mo agoConsidering these are essentially document completion engines[0], can't you just start the task with the version that doesn't refuse and then continue with the version that would refuse but now has to keep going after it accepted the task? :-P [0] in the sense that the "discussion" is basically a turn based game between you and the LLM filling a chat transcript document
- edg5000 1mo agoThe article proves your point. It's eventually proceeded once there was the right history. If he already had the exploit and wanted the model to write an implementation, faking the history could have tripped the model over the edge. So the guardrails aren't unsurmountable.
- DiabloD3 1mo agoIt does depending on the technique.
- miroljub 1mo agoA bit worse quality is a fine trade off when the alternative is no output (zero quality).
- timmmmmmay 1mo agoEarly attempts at this sort of thing definitely did, but these days the impact is minimal
- Aurornis 1mo ago> There are versions of Qwen3.8-27B that are unrestricted and available from hugging face. The restrictions are not a single check in the model that can be removed. Those models on Huggingface are manipulated in different ways that also degrade the model’s intelligence. The degradation ranges from subtle to obviously broken, but it’s not free. When the restrictions are built into the model’s training sets you can try to alter the weights that are involved in the refusals, but that doesn’t mean that what’s left is useful or good knowledge for the same task. Those weights also might be involved in other tasks, so altering them can interfere with interactions that aren’t obviously related.
- rustcleaner 1mo ago>There are versions of Qwen3.8-27B that are unrestricted and available from hugging face. Based! :DDD The uncensorers are oblique, if not parallel, to machine learning Robin Hoods. May their efforts continue indefinitely, or at least until the likes of Altman and Amodei are bankrupt and crying into their low fat Cherios!