4 ms·
I wonder if llm are biased towards older, more insecure implementations because there is a higher volume of old code vs new code. Same thing with the data it i
by pylua 3y ago
I wonder if llm are biased towards older, more insecure implementations because there is a higher volume of old code vs new code.
Same thing with the data it is trained on — not all code requires all levels of refinement. Most of the data is probably around average.
- squigz 3y agoI'm not sure that code being newer inherently means it will be more secure
- pylua 3y agoI don’t think it is a tautology , but I can imagine a cve scanner picking up older code with log4j where newer code may avoid that library altogether, just as an example. Since there is more older code than newer code would the llm be suspectible to that ?
- xmcqdpt2 3y agoEveryone still uses log4j after the fix. There aren't many full-featured alternatives... and the one that exist probably contain unfixed bugs.
- darkerside 3y ago> Most of the data is probably around average. I know this is not how distributions work, but I had to chuckle at the literal interpretation of this.
- d-z-m 3y agoI'd say the data is pretty normal.
- foota 3y agoThis makes me wonder about training an LLM on one language and then fine tuning it for another. If you train over only, say, JavaScript, and then finetune for C, I imagine it will be quite bad at writing safe code, even if it makes the code look like C, because it didn't have to learn about freeing and such. Similarly, would it pick up patterns from one language and keep then in the other? Maybe an LLM trained on Kotlin would be more likely to write functional code finetuned.
- fragmede 3y agoGiven that dataset anomalies can result in LLM output corruption, I'm not convinced that cross-training like that would even work.