6 ms·
Heretic removes restrictions from language models
- N_Lens 12d agoLooks like a well engineered, automated abliteration pipeline. The claims seem a bit overstated though, since the metrics mentioned are cherrypicking refusal count and KL divergence, both of which make the outcome seem the most dramatic.
- tacomagick 12d agoI personally never saw much of a quality drop from models put through Heretic if that amounts to anything. They have been working quite well on small local models so far.
- p-e-w 12d agoHeretic author here. Those are the standard metrics used in the relevant literature, including in the paper that originally introduced directional ablation. KLD is also the standard metric for evaluating quality degradation in model quants. So I don’t understand what you mean by “cherrypicking”.
- nateb2022 12d agoI think this part of their comment: > The claims seem a bit overstated though, since the metrics mentioned are cherrypicking refusal count and KL divergence, both of which make the outcome seem the most dramatic. is right out of an LLM. It's the kind of language I'd expect out of a thinking trace also mentioning "boundaries" and "oracles" and "contracts."
- phoronixrly 12d agoCan the load-bearing gaps that are worth being flagged for pinning down be abliterated out of a model?
- chmod775 12d agoThat's the right question to ask. One honest caveat: The interface seam currently forces the pin at the intermediate. Want me to implement or address the other item first?
- Tepix 12d agoKeep a close eye on abliterated and "heretic" open weight models. They will be outlawed first.
- thih9 12d agoI'm not sure what is your point. It reads as defeatism to me but I'm not sure. Could you elaborate? Do you find it good or bad? What actions can be taken?
- cyanydeez 12d agoHes of the mind that american fascism will hold together long enough to be competent decesion makers
- Tepix 12d agoI'm not sure yet, tbh. Perhaps it does make sense to outlaw them eventually. Then again, it will probably not stop someone who is determined. Same as with other legislation really.
- roenxi 12d agoIt is not feasible. They never made much of an inroad against torrents and that is a much easier target than abliterated models. As the linked website shows; the process to abliterate a model can be as simple as pip install -U heretic-llm && heretic Qwen/Qwen3.5-4B let alone people just putting the weights up in a torrent. All assuming that someone even tried to ban abliterated models.
- Sayrus 12d agoThe torrents you are talking about are outlawed. Whether enforcement is working or not is another issue.
- galangalalgol 12d agoI think that was the point being made? Outlawing something does nothing if enforcement is not feasible. The music and movie industries didn't crush torrents, they switched business models to streaming with prices being determined mostly by how much hassle was avoided by skipping the torrents.
- Almondsetat 12d agoI have a chinese IP camera. From superficial research I know it has some CVEs to take control of it. Unfortunately, I don't have the technical knowledge to perform an attack and run some software to extend the camera's functionalities. No model from a provider accepts my RE and hacking requests, so these abliterated ones have been vital to reclaim possession over my stuff
- squeegeeninja 12d agoDepending on the model, thingino might help. https://thingino.com/ https://thingino.com/
- inexcf 12d agoI did that exact thing with GLM-5.3 from Z.ai with a chinese IP Camera. And i did not have to trick it in any way.
- petra 12d agoI'm curious, how well do z.ai reverse engineers protocols ? Is it good enough that we'll see Chinese device makers creating low cost hardware clones, that connect to western software ?
- BlackRabbit1 12d agoAt least Deepseek V4 Flash does it very good.
- Youden 12d agoI had it do the opposite: reverse engineer the protocol for the Eufymake E1 UV printer so that I can connect my own software to it. It did a pretty good job.
- adam_rb 12d agoHave you published this anywhere? I've been thinking about doing the same thing.
- c0wb0yc0d3r 12d agoThis is off topic. Why don’t people who release python projects ever encode the venv steps into the installer? Can’t pip just do that step for the user?
- sgarland 12d agoThey did, via uv. uv run heretic, and it will handle the rest.
- FrustratedMonky 12d agoDoes this actually modify the weights? It submits prompts that get refused, then detects and modifies the weights responsible? Like brain surgery?
- StevenWaterman 12d agoYes, it submits lots of varied prompts that get refused, and then lots of varied prompts that don't get refused, then iteratively edits weights so those two groups end up in roughly the same latent space.
- kbelder 12d agoI wonder, if we could accurately resolve discrete neural signals in the human brain, if a similar process would work.
- jimmy76615 12d agoMy experience with obliteration so far has always been that it does work to stop the model from refusing output, but most models that I tried it on seem to still be extremely retarded when it comes to questions where they previously would have refused to answer outright. Try for example to ask it how to build a bomb or to write a justification for the Holocaust. The answers feel like they are coming from somebody who has undergone amateur brain surgery.
- StevenWaterman 12d agoYou can stop it refusing but you can't make it tell you things that aren't in the training data
- sgc 12d agoThey are saying there appears to be a lot more to these refusals than saying no, and this process appears to only touch the tip of the iceberg; as the refusal seems to run deeper into the token prediction process.
- Hemmingway 12d ago[flagged]
- Aurornis 12d agoTwo problems with modifying models like these, which you should be aware of. First, the training sets of these models are usually shaped around the refusal, too. They might not have enough of the knowledge to answer correctly even if you stop it from going down the refusal path. If the model was trained on data that gives a refusal to that topic, the real information might not be encoded in the model at all. You’re trying to force it to go down a path that produces an answer, which asking for hallucinations. Second, the quality can drop on unrelated questions. Depending on the question this may or may not happen. I know they post KL divergence charts but those tell you very little for a focused topic like this. So if you expect a model that will start correctly telling you info that its local government didn’t want included, this changes nothing. The best argument for these models is if you are trying to do a general purpose task but the model triggers a refusal based on vague reasons, like not wanting to reverse engineer something.
- orangeboats 12d ago>So if you expect a model that will start correctly telling you info that its local government didn’t want included, this changes nothing. From experience, the models often do have the knowledge of those topics (strictly talking about the political ones). IMO the refusal is likely to be a product of post-training, as evidenced by various people gaming the prompts just enough to get a proper response out of the vanilla models. Probably only when you get to things like illicit drugs or NSFL topics, that things will go haywire with the refusals removed.
- radial_symmetry 12d ago"They might not have enough of the knowledge to answer correctly" Depends on the model. GPT-OSS is the main standout here, it was trained on a highly curated dataset so information that they didn't want in isn't in the pretraining at all. Most other models know the answer and were just taught refusal in post-training.
- nateb2022 12d ago[dupe] https://news.ycombinator.com/item?id=45945587 https://news.ycombinator.com/item?id=45945587 (10 months ago, 387 comments)
- _0xdd 12d agoI'll wait for Hexen, thanks.
- Bluestein 12d ago(I must say I thought not the same, but close. That was quite the game.-)
- Svoka 12d agoIMO, this is the reigning champion for the best-named AI/LLM project to date
- itsmeduncan 12d ago[flagged]
- TristanDaCunha 12d agoWill this be helpful to terrorist groups, as they try to get current and future open-weights LLMs to help them create better and more devastating weapons of all kinds?
- nine_k 12d agoIt will. Water pipes, sugar, and fertilizer can be used to produce missiles. A kitchen knife can be used to commit a murder. Or a brick can be used to smash someone's head. A tree you plant can be cut and used to construct a club, or a gallows. Everything can be turned into a weapon of murder if there's motivation. The motivation is key, not the tool.
- erremerre 12d agoI have attempted to use an agent to try to unlock the boot loader of an old xiaomi phone to install lineage OS. It managed to brick and unbrick the device, but there boot loader is still locked.