4 ms·
We sorely need a way to reliably detect AI slop, but unfortunately it doesn't seem possible and it's just getting harder and harder. Last month I tried my hand
by pscanf 5mo ago
We sorely need a way to reliably detect AI slop, but unfortunately it doesn't seem possible and it's just getting harder and harder.
Last month I tried my hand at finding a way to tell whether an OSS project is slop or not, based on the amount of "human attention" it received vs the amount of code it contains. The idea is that a 100k LOC project which received 3 days' worth of attention from a human is most certainly slop.
The approach doesn't work very well, though¹, mostly because it's hard to gauge the amount of attention that was given. If I see one commit with +3000 LOC, I can assume it's AI-generated, but maybe you're just the type of dev that commits infrequently.
Maybe we need some sort of "proof of human attention" for digital artifacts, that guarantees that a human spent X time working on it.
¹ I wrote about it here https://pscanf.com/s/352/ https://pscanf.com/s/352/
- ChrisMarshallNY 5mo agoI suspect that it will be impossible, soon. People will just train LLMs to "act human," and pass the various turing tests we throw at them. I stay pretty busy[0], and have been accused of "gaming" my GH repos. That's not the case. I'm retired, experienced, and working on software all day, every day. I just don't get paid for it. I also don't especially care, whether or not anyone thinks I'm a bot. I eat my own dogfood. Most of my work is on modules that I use in my own projects. [0] https://github.com/ChrisMarshallNY#github-stuff https://github.com/ChrisMarshallNY#github-stuff
- caymanjim 5mo agoThere's no reason to care that a human spent time on it. Humans are bad at writing code. Garbage PRs and slop have been a problem in open source and bug bounty programs since long before AI came on the scene. We need better AI so that there's no need to solicit external bug fixes, and better AI so other contributions can be evaluated for usefulness and quality. What do you care if a human ever looked at it at all? It implies that humans are adding value to the process. It's possible for a human to add value. The right human can add tremendous value. But I'll take a completely autonomous AI over 99% of the human software engineers and 99% of the people contributing PRs and bugfixes. It was hard to keep up with slop before. It's a lot harder now. AI will help weed through the garbage.
- bcjdjsndon 5mo agoReasonable logic but I bet you get downvoted
- 48terry 5mo agoIf AI is already mass-producing garbage PRs and other unreliable crap, what makes AI (established as producing unreliable crap) the solution for review? What makes the reviewing AI not produce unreliable crap with regards to the review? A magical, hypothetical AI that always gets it right and will make all these problems go away is neither a solution nor a plan. It's wishful thinking.
- caymanjim 5mo agoAI in the hands of the right people is incredibly powerful. A good team of engineers with AI doing their own bug-hunting on their own code is already far better than any outsider—human, AI, or human-assisted AI—could ever do. A good internal AI-assisted team is also the only thing that can vet all other contributions. It doesn't matter if those contributions are 100% human-written, 100% AI-written, or a combination. The problem is the same. Unless you stop accepting outside contributions at all, there's simply no way to determine if a human was involved in the process. Any mandate that all contributions come from humans will fail because there's no detection or enforcement mechanism. You have to assume it's slop either way, and improve your ability to vet it. Only another AI can do that, because we don't have enough qualified humans to keep up.
- 48terry 5mo agoThat didn't actually address my comment or question, so I'll repeat it, I guess. We already know AI is spamming unreliable crap and slop. The apparent solution is "more, better AI". Why wouldn't this AI for screening all this also produce crap and slop? Is the plan there "AI but it actually works right and doesn't produce crap and slop"?
- caymanjim 5mo ago
- theragra 5mo agoYou by default assume all AI code is slop? This seems like an approach that won't be beneficial to the OSS overall. What I found is 1. With LLMs, I was finally able to find time and confidence to contribute a patch. Before, contributing to established project seemed impossible with an amount of guidelines to follow and insider knowledge to have. 2. Many small isolated parts do not require some great code. Just glue, tying libraries together allows producing new features in existing apps. With some cars, it is possible even if I don't know the languages and libraries used.