7 ms·
You're not just using a tool — you're co-authoring the science. This README is an absolute headache that is filled with AI writing, terminology that doesn't ex
by a2128 7mo ago
You're not just using a tool — you're co-authoring the science.
This README is an absolute headache that is filled with AI writing, terminology that doesn't exist or is being used improperly, and unsound ideas. For example, it focuses a lot on doing "ablation studies", by which it means removing random layers of an already-trained model, to find the source of the refusals(?), which is an absolute fool's errand because such behavior is trained into the model as a whole and would not be found in any particular layer. I can only assume somebody vibe-coded this and spent way too much time being told "You're absolutely right!" bouncing back the worst ideas
- creatonez 7mo ago> For example, it focuses a lot on doing "ablation studies", by which it means removing random layers of an already-trained model, to find the source of the refusals(?), which is an absolute fool's errand because such behavior is trained into the model as a whole and would not be found in any particular layer. That doesn't mean there couldn't be a "concept neuron" that is doing the vast majority of heavy lifting for content refusal, though.
- mapontosevenths 7mo agoThats not what it means at all. It uses SVD[0] to map the subspace in which the refusal happens. Its all pretty standard stuff with some hype on top to make it an interesting read. Its basically using a compression technique to figure out which logits are the relevant ones and then zeroing them. [0] https://en.wikipedia.org/wiki/Singular_value_decomposition https://en.wikipedia.org/wiki/Singular_value_decomposition
- D-Machine 7mo agoYou are also not quite correct, IMO. See my comment at https://news.ycombinator.com/item?id=47283197 https://news.ycombinator.com/item?id=47283197. What you are talking about is abliteration. What OBLITERATUS seems to be claiming to do is much more dumb, i.e. just zeroing out huge components (e.g. embedding dimension ranges, feed-forward blocks; https://github.com/elder-plinius/OBLITERATUS?tab=readme-ov-file#ablation-strategies https://github.com/elder-plinius/OBLITERATUS?tab=readme-ov-f...) of the network as an "Ablation Study" to attempt to determine the semantics of these components. However, all these methods are marked as "Novel", I.e., maybe just BS made up by the author. IMO I don't see how they can work based on how they are named, they are way too dumb and clunky. But proper abliteration like you mentioned can definitely work.
- mapontosevenths 7mo agoYou got me there. I missed the wackier antics further down. Mea culpa.
- D-Machine 7mo agoSo did I initially until I saw a few more things from others here.
- dinunnob 7mo agoHmm, pliny is amazing - if you kept up with him on social media you’d maybe like him https://x.com/elder_plinius https://x.com/elder_plinius
- bigyabai 7mo agoIf this qualifies as "amazing" in 2026 then Karpathy and Gerganov must be halfway to godhood by now.
- dinunnob 7mo agoI dont think anyone is going to dispute this
- bigyabai 7mo agoI just don't think many people will be "amazed" by their output, as you claim.
- dinunnob 7mo agoI just said pliny was amazing, fwiw - i like that hes hacking on these and posts about it. I rushed to defend, i wish more people were taking old school anarchist cookbook approaches to these things
- robertk 7mo agoYou don't know what you are talking about. Obviously refusal circuitry does not live in one layer, but the repo is built on a paper with sound foundations from an Anthropic scholar working with a DeepMind interpretability mentor: https://scholar.google.com/citations?view_op=view_citation&hl=en&user=NgyIgX4AAAAJ&citation_for_view=NgyIgX4AAAAJ:qjMakFHDy7sC https://scholar.google.com/citations?view_op=view_citation&h...
- paradox460 7mo agoIt's not just a headache, it's bad
- Retr0id 7mo agoI don't know if this particular tool/approach is legit, but LLM ablation is definitely a thing: https://arxiv.org/abs/2512.13655 https://arxiv.org/abs/2512.13655
- D-Machine 7mo agoDoesn't look legit to me. You are talking about abliteration, which is real. But the OP linked tool is doing novel and very dumb ablation: zeroing out huge components of the network, or zeroing out isolated components in a way that indicates extreme ignorance of the basic math involved. Compared to abliteration, none of the ablation approaches of this tool make even half a whit of sense if you understand even the most basic aspects of an e.g. Transformer LLM architecture, so my guess is this is BS.
- hexaga 7mo agoThe terminology comes from the post[0] which kicked off interest in orthogonalizing weights w.r.t. a refusal direction in the first place. That is, abliteration was not originally called abliteration, but refusal ablation. Ultimately though, OP is just what you get if you take the idea of abliteration and tell an LLM to fix the core problems: that refusal isn't actually always exactly a rank-1 subspace, nor the same throughout the net, nor nicely isolated to one layer/module, that it damages capabilities, and so on. The model looks at that list and applies typical AI one-off 'workarounds' to each problem in turn while hyping up the prompter, and you get this slop pile. [0]: https://www.lesswrong.com/posts/refusal-in-llms-is-mediated-by-a-single-direction#Feature_ablation_via_weight_orthogonalization https://www.lesswrong.com/posts/refusal-in-llms-is-mediated-...
- jandrese 7mo agoNo offense, but a Lesswrong link is an immediate yellow flag, especially on the topic of AI. I can’t say if that article in particular is bad, but it is associating with a whole lot of abject nonsense written by people who get high on their own farts.
- fragmede 7mo agoAlternately, it's intentional. It very effective filters out people with your mindset. You can decide if that's a good thing or not.
- eli 7mo agoWhy would a tool that works need to dissuade skeptics from trying it?
- dmix 7mo agoBased on his twitter he may just like irony/meta posting a little too much like a lot of modern culture
- D-Machine 7mo agoI immediately read it as intentional, as a sort of attempt at ironic / nihilistic humour re: LLM-generation, given what the tool claims to do.
- deleted 7mo ago[deleted]
- lazzlazzlazz 7mo ago[flagged]
- SV_BubbleTime 7mo agoIt doesn’t even surprise me anymore. The people here think they’re so superior to the already arrogant redditors… same people. Thing definitely exists… some top level comment somewhere telling about how it doesn’t exist.
- lazzlazzlazz 7mo agoExactly. And I'm downvoted below 0 for pointing this out. :)
- deleted 7mo ago[deleted]
- D-Machine 7mo ago> "ablation studies", by which it means removing random layers of an already-trained model, to find the source of the refusals(?) This is not what an ablation study is. An ablation study removes and/or swaps out ("ablates") different components of an architecture (be it a layer or set of layers, all activation functions, backbone, some fixed processing step, or any other component or set of components) and/or in some cases other aspects of training (perhaps a unique / different loss function, perhaps a specialized pre-training or fine-tuning step, etc) in order to attempt to better understand which component(s) of some novel approach is/are actually responsible for any observed improvements. It is a very broad research term of art. That being said, the "Ablation Strategies" [1] the repo uses, and doing a Ctrl+F for "ablation" in the README does not fill me with confidence that the kind of ablation being done here is really achieving what the author claims. All the "ablation" techniques seem "Novel" in his table [2], i.e. they are unpublished / maybe not publicly or carefully tested, and could easily not work at all. From later tables, I am not convinced I would want to use these ablations, as they ablate rather huge portions of the models, and so probably do result in massively broken models (as some commenters have noted in this thread elsewhere). EDIT: Also, in other cases [1], they ablate (zero out) architecture components in a way that just seems incredibly braindead if you have even a basic understanding of the linear algebra and dependencies between components of a transformer LLM. There is nothing sound clearly about this, in contrast to e.g. abliteration [3]. [1] hhtps://github.com/elder-plinius/OBLITERATUS?tab=readme-ov-file#ablation-strategies [2] https://github.com/elder-plinius/OBLITERATUS?tab=readme-ov-file#novel-techniques-2025-2026 https://github.com/elder-plinius/OBLITERATUS?tab=readme-ov-f... EDIT: As another user mentions, "ablation" has a specific additional narrower meaning in some refusal analyses or when looking at making guardrails / changing response vectors and such. It is just a specific kind of ablation, and really should actually be called "abliteration", not "ablation" [3]. [3] https://huggingface.co/blog/mlabonne/abliteration https://huggingface.co/blog/mlabonne/abliteration, https://arxiv.org/abs/2512.13655 https://arxiv.org/abs/2512.13655.
- hexaga 7mo agoWhat do you mean? It's a spin on abliteration / refusal ablation. Roughly, from what I remember abliteration is: 1. find a direction corresponding to refusal by analyzing activations at various parts of a model (iirc, via mass means seen earlier in Marks, Tegmark and shown to work well for similar tasks) 2. find the best part(s) of the model to orthogonalize w.r.t. that direction and do so (exhaustive search w/ some kind of benchmark) OP is swapping in SVD for mass means (1), and the 'ablation study' for (2), and a bunch of extra LLM slop for... various reasons. The final model doesn't have zeroed chunks, that is search for which parts to orthogonalize/refusal ablate/abliterate. I don't have confidence that it works very well either, but, it isn't 'braindead' / obvious garbage in the way you're describing. It's LLMified but standard abliteration. The idea has fundamental limitations and LLMs tend to work sideways at it -- there's not much progress to be made without rethinking it all -- but it's very conceptually and computationally simple and thus attractive to AIposters. You can see how the LLMs all come up with the same repackaged ideas: SVD does something deeply similar to mass means (and yet isn't exactly equivalent, so LLM will _always_ suggest it), the various heuristic search strategies are competing against plain exhaustive search (which is... exhaustive already), and any time you work with tensors the LLM will suggest clipping/norms/smoothing of N flavors "just to be safe". And each of those ends up listed as "Novel" when it's just defensive null checks translated to pytorch. I mean, the whole 'distributed search' thing is just because of how many combinations of individual AI slops need to be tested to actually run an eval on this. But the idea is sound! It's just terrible. I'm not defending the project itself -- I think it's a mess of AIisms of negligible value -- but please at least condemn it w.r.t. what is actually wrong and not 'on vibes'.
- jeffbee 7mo ago"Ablation studies" are a real thing in LLM development, but in this context it serves as a shibboleth by which members of the group of people who believe that models are "woke" can identify each other. In their discourse it serves a similar purpose to the phrase "gain of function" among COVID-19 cranks. It is borrowed from relevant technical jargon, but is used as a signal.
- 06867457397658 7mo ago[flagged]
- gopher_space 7mo agoPositive keywords in this area of interest would be "point of view", "subtext", and "Art Linkletter".
- drnick1 7mo agoI wouldn't call mainstream LLMs "woke," but they are definitely on the "politically correct" side of things. There should be NO restriction on open source models. They should just reflect the state of human knowledge and not take a stance on whether some activity is illegal or immoral.
- pjc50 7mo agoDefining morality out of the set of knowledge is quite an opinion.
- simondotau 7mo agoA model should understand multiple perspectives on morality and avoid prescribing a single one where there’s no overwhelming prior consensus. Alternatively, they should be trained on my opinion on everything. That would also be acceptable.
- simgt 7mo agoIf LLMs were a public good released by non profit entities, that could make sense, maybe. Turns out spewing illegal and immoral shit is not good for the PR of most for-profit businesses.
- userbinator 7mo ago"Getting high on your own supply" is exactly what I'd expect from those immersed in this new AI stuff.
- shevy-java 7mo agoIs that quote from the movie Scarface? https://www.youtube.com/watch?v=U4XplzBpOiU https://www.youtube.com/watch?v=U4XplzBpOiU # had to search for it right now, seems to be a movie-quote \o/
- DeathArrow 7mo ago> I can only assume somebody vibe-coded this and spent way too much time being told "You're absolutely right!" bouncing back the worst ideas Are there LLMs which don't always approve whatever idea the user has and tell him it's absolutely brilliant?
- isjdiwjdus 7mo ago[dead]