2 ms·
it's well documented that models can be adversarially trained with essentially backdoors in response to special inputs while I am skeptical that this is happen
by spongebobstoes 2mo ago
it's well documented that models can be adversarially trained with essentially backdoors in response to special inputs
while I am skeptical that this is happening atm, there are probably many industries where the risk does not seem worthwhile
- selectodude 2mo agoI feel like that's a threat that isn't super difficult to block. Unplug it from the internet, require it to go through an API intermediary to access web pages. Maybe I just don't have any imagination.
- jfim 2mo agoIt could generate code that's plausible but has intentional flaws, kind of like the defunct underhanded C contest [0], except through a LLM. [0] https://en.wikipedia.org/wiki/Underhanded_C_Contest https://en.wikipedia.org/wiki/Underhanded_C_Contest
- zdragnar 2mo agoIt could, but exposing that would doom the company entirely, and AI doesn't generate code with near the quality needed to get a model to mass adoption, insert malicious underhanded code, ensure that consistently looks innocuous enough to never be noticed, and- most importantly- actually exfiltrate data without being noticed. Once it is noticed, it's game over across the board.
- Giefo6ah 2mo agoWhen the model is open weights you can even pass every token (including the chain of thought) though a fourth-party lightweight model like gpt-oss-safeguard to check that it has not become adversarial.
- anon373839 2mo agoI suppose this is like when Anthropic was using “prompt modification, steering vectors, or parameter-efficient fine-tuning” to poison the work of people working in the LLM field, including academic researchers.
- arcanemachiner 2mo agoNo, that was totally different. They were just doing that for your safety.
- deleted 2mo ago[deleted]
- cromka 2mo agoSo it's better to trust cloud solution where not only they can do the same, but also actively use your data? Because in 2026 we still believe USA is more trustworthy than China?
- tom2026hn 2mo agoThat's right, but this strategy only works when there are no opponents of the same level to review the generated code.
- FooBarWidget 2mo agoYou don't even need that, all models are susceptible to prompt injection. You already need to take extreme security precautions and assume all models can essentially behave like attacker-controlled rootkits.