3 ms·
They've said that they'll stop notifying developers when this gets triggered, instead they'll load in basically like a LORA that's designed to inject bugs into
by mips_avatar 4mo ago
They've said that they'll stop notifying developers when this gets triggered, instead they'll load in basically like a LORA that's designed to inject bugs into your code.
- HDBaseT 4mo agoAntrophic wants to stop training models and ride out Mythos / Fable for as long as possible. They are trying to expand the 6-18 month gap they have against China-based models. Could the gap widen to say 24 months behind?
- p-e-w 4mo agoTheir gap over Chinese models like GLM-5.1 is nowhere near 18 months. In many areas, it’s less than 6 months. The best closed models 18 months ago were worse than Qwen3.6.
- echelon 4mo agoThese coding agent models only started getting useful in January. Before that they were difficult to control autocomplete, and not very smart. January was an inflection point, and no open weights model has crossed over that same threshold. This is definitely recursive self improvement territory, except that we're prohibited from participating. It feels like the capability gap is wider than before.
- slopinthebag 4mo agoIt was more like November. But it wasn’t really an inflection point, harnesses got good enough that people started noticing by the holiday break. And I’m not discounting some good ol’ stealth marketing in there as well. Deepseek feels pretty close to Opus at this point, and it’s certainly useful enough for me to spend $20 on api tokens instead of four Claude max plans….
- lbreakjai 4mo agoHave you tried deepseek V4? It costs pennies and is as good as Opus 4.6 (I found 4.7 to be a downgrade, and cancelled my claude subscription before 4.8). The threshold has definitely been crossed.
- echelon 4mo agoIt is not as good as Opus. I've tried to write Rust with it (and Codex for that matter), and it's awful.
- nomel 4mo ago> a LORA that's designed to inject bugs into your code A statement like this, clearly, requires a reference.
- mips_avatar 4mo agoFrom the model card: "the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning" aka they will take your ML research code and inject bugs into it until it breaks using a LORA (or some other form of PEFT)
- nomel 4mo agoThanks, I thought maybe I missed something. That's an interesting way to interpret that.
- giancarlostoro 4mo agoPEFT is a library, one of its capabilities is to produce LoRAs. See: https://heidloff.net/article/efficient-fine-tuning-lora/ https://heidloff.net/article/efficient-fine-tuning-lora/
- adw 4mo agoIt's just an acronym, "parameter-efficient fine tuning". LoRA is one method, prefix tuning is another, there are more.
- mips_avatar 4mo agoAnthropic is trying to hide bad behavior by being vague, it's important to not be vague when calling it out.
- nomel 4mo agoI'm of the opinion that removing guardrails is how you force regulation. What's your opinion on the balance?