6 ms·
Establishing an etiquette for LLM use on Libera.Chat
- aspenmayer 2y agoHN would benefit from a specific, explicit policy such as this.
- benatkin 2y agoNope. > LLMs are allowed on Libera.Chat. They may both take input from Libera.Chat and output responses to Libera.Chat. This wouldn't help HN. Nor would the opposite policy, if only because it would encourage accusatory behavior.
- t-writescode 2y agoThe odds of LLMs being used to produce content on HN is a number approaching 100%. The odds of LLMs being trained / queried against data scraped from HN or HNSearch is even closer to 100%. I know you don't like the "LLMs are allowed..." part, but they're here and they literally cannot be gotten rid of. However, this rule, > As soon as possible, people should be made aware if they are interacting with, or their activity is being seen by, a LLM. Consider using line prefixes, channel topics, or channel entry messages. Could be something that is strongly encouraged and helpful, and possibly the "good" LLM users would follow it.
- aspenmayer 2y agoI have asked dang to comment on this issue specifically in the context of this post/thread. The “opposite policy” is sort of the current status quo, per dang: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=by%3Adang%20%22generated%20comments%22&sort=byDate&type=comment https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... See this thread for my own reasoning on the issue (as well as dang’s), as it was raised recently: https://news.ycombinator.com/item?id=41937993 https://news.ycombinator.com/item?id=41937993 You’ll need showdead enabled on your profile to see the whole thread, which speaks to the controversial nature of this issue on HN. I agree that your mention of “encouraging accusatory behavior” is a point well-taken, and in the absence of evidence, such accusations themselves would likely run afoul of the Guidelines, but it’s worth noting that dang has said that LLM output itself is generally against the Guidelines, which could lead to a feedback loop of disinterested parties posting LLM content, only to be confronted with interested parties posting uninteresting takedowns of said LLM content and posters of it. No easy answers here, I’m afraid.
- benatkin 2y agoFrom the thread with see this thread > There are lot of grey areas; for example, your GP comment wasn't just generated—it came with an annotation that you're a lawyer and thought it was sound. That's better than a completely pasted comment. But it was probably still on the wrong side of the line. We want comments that the commenters actually write, as part of curious human conversation. This doesn't leave much room for AI non-slop: > We want comments that the commenters actually write, as part of curious human conversation. I think HN is trying to be good at being HN, not just to provide the most utility to its users in general. So those wanting something like HN if it started in 2030, may want to try and build a new site.
- refulgentis 2y agoLaw is hard! In general, the de facto status quo is: 1. For whatever reason*, large swaths of LLM output copy-pasted is easily detectable. 2. If you're restrained, polite, with an accurate signal on this, you can indicate you see this, and won't get downvoted heavily. (ex. I'll post "my internal GPT detector went off, [1-2 sentence clipped version of why I think its wrong even if it wasn't GPT]") 3. People tend to downvote said content, as an ersatz vote. In general, I don't think there needs to be a blanket ban against it, in the sense of I have absolutely no problem with LLM output per se, just lazy invocation of it, ex. large entry-level arguments that were copy-pasted. i.e. I've used an LLM to sharpen my already-written rushed poor example, which didn't result in low-perplexity, standard-essay-formatted, content. Additionally, IMHO it's not bad, per se, if someone invests in replying to an LLM. The fact they are replying indicates its an argument worth furthering with their own contribution. * a strong indicator that a fundamental goal other than perplexity minimization may increase perceived quality
- famouswaffles 2y agoThe reason is not strange or unknown. The text completion GPT-3 from 2020 often sounds more natural than 4. The reason is the post training processes. Models are more or less being trained to sound like that during RLHF. Stilted, robotic, like a good little assistant. Open AI, Anthropic have said as much. It's not a limitation of the loss function or even state of the art.
- jsheard 2y agoThe community already seems to have established a policy that copy pasting a block of LLM text into a comment will get you downvoted into oblivion immediately.
- aspenmayer 2y agoThat rubric only works until sufficiently advanced LLM-generated HN posts are indistinguishable from human-generated HN posts. It also doesn’t speak to the permission or lack thereof of training LLMs on HN content, which was another main point of OP.
- deleted 2y ago[deleted]
- JavierFlores09 2y ago> That rubric only works until sufficiently advanced LLM-generated HN posts are indistinguishable from human-generated HN posts. if a comment made by a LLM is indistinguishable from a normal one, it'd be impossible to moderate anyway unless one starts tracking people across comments and see the consistency of their replies and overall stance so I don't particularly think it is useful to worry about people who will go the extra length to go undetected
- aspenmayer 2y ago> if a comment made by a LLM is indistinguishable from a normal one, it'd be impossible to moderate anyway unless one starts tracking people across comments and see the consistency of their replies and overall stance so I don't particularly think it is useful to worry about people who will go the extra length to go undetected The existence of rule-breakers is not itself an argument against a rules-based order.
- tredre3 2y agoHN's guidelines aren't "laws" to be "enforced", they're a list of unwelcome behaviors. There is value in setting expectations for participants in a community, even if some will choose to break them and get away with it.
- tptacek 2y agoWe have an explicit policy: you can't post LLM stuff directly to HN. https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=human%20generated%20by%3Adang&sort=byDate&type=comment https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
- aspenmayer 2y agoThis policy should manifest itself in the Guidelines, if HN users are expected to know about it and adhere to it.
- Jerrrry 2y agoIts human users can infer it, the other uses can't, yet.
- throwaway314155 2y agoDoesn't seem very explicit if you have to search a mod's comment history to find it.
- tptacek 2y agoLots of rules on HN work that way. It's a whole thing. We probably don't need to get into it here. I think it works pretty well as a system. We have a jurisprudence!
- Jerrrry 2y agoWell said, "Better unsaid." Shame is the best moderator. Also, HN's miscellaneous audience of rule breakers benefit from having some rules be better off not stated. Especially this one, as it is almost as good as a "Gun-Free Zone"
- fngjdflmdflg 2y agoI don't think that is correct. Dang usually links directly to the guidelines and even quotes the exact guidelines being infringed upon sometimes. '"dang" "newsguidelines.html"' returns 20,909 results on algolia.[0] (Granted, not all of these are by Dang himself, I don't think you can search by user on algolia?) Some of the finer points relating to specific guidelines may no be directly written there, eg. what exactly is considered link bait or not etc., but I don't think there are any full blown rules not in the guidelines. I think the reason LLMs haven't been added is because it's a new problem and making a new rule to quickly that may have to change later will just cause more confusion. [0] https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=%22dang%22%20%22newsguidelines.html%22&sort=byDate&type=comment https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
- superkuh 2y agoMostly it's just formalizing of the established status quo. But the changes re: allowing training on chat logs has caused some unintended consequences. For one, now the classic IRC megahal bots which have been around for decades are technically not allowed unless you get permission from Libera staff (and the channel ops). They are markov chains that continuously train on channel contents as they operate. But hopefully, as in the past, the Libera staffers will intelligently enforce the spirit of the rules and avoid any silly situations like the above caused by imprecise language.
- comex 2y agoBy its wording, the policy is specifically about training LLMs. A classic Markov chain may be a language model, but it’s not a large language model. The same rules might not apply.
- superkuh 2y agoYeah, you'd think, but this one was run by the staff in #libera the other night after the announcement and it sounded like they believed markovs technically counted. But I imagine as long as no one is rocking the boat they'll be left alone. Perhaps there was some misunderstanding on my part.
- martin-t 2y agoA classic example of a community self regulating until overwhelmed at which point rules are imposed which bad previously accepted and harmless behavior. Rules must take scale into account and do it explicitly to avoid selective enforcement. There's a difference between one person writing a simple bot and a large corporation offering a bot pretending to be human to everyone. The first is harmless and fun, the second is a large scale for-profit behavior with proportionally large negative externalities.
- fjdjshsh 2y agoI strongly believe it should be illegal to post something automatically by an LLM without clearly identifying it as such. I hope countries start passing these laws soon
- conception 2y agoWhy pass a law that’s completely unenforceable?
- theamk 2y agoIt is somewhat enforceable. Sure, no one is going to go after random reddit post, but if a Major Newspaper wants to have AI write their articles, this would have to be labeled. And if your bank gets LLM support agent, it can no longer pretend to be human. All very desireable outcomes IMHO.
- karlgkk 2y agoIt's not unenforceable at all. Major players would be forced to abide by it, smaller players would reduce their use of LLMs, and not marking LLM content would be a bannable offense on most platforms.
- conception 2y agoOpen source Local LLMs are already a thing. That pandora’s box is way way open already.
- Grimblewald 2y agoblack market guns are also a thing, and they're realtivly easy to manufacture using untracked machines and materials using skills you can develop in under a year. That doesn't mean that regulating the sale / supply / ownership of guns isn't useful.
- Jensson 2y agoIf a big company does that they wouldn't be able to hide it, this is just as easy to enforce as any other regulation.
- ranger_danger 2y agoNow can libera please establish etiquette for channel mods? All the biggest channels have extremely toxic, egotistical mods with god complexes visible from space.
- aspenmayer 2y agoHave you seen examples of such codes of conduct in the IRC context before? Closest thing I can think of maybe is SDF’s or other shared systems’, but such rules seem somewhat quaint compared to norms on IRC. Speaking of SDF, here’s their bot policy: https://sdf.org/?faq?CHAT?01 https://sdf.org/?faq?CHAT?01 > [01] CAN I RUN AN IRC BOT HERE?? > IRC BOTs are pretty intensive and most systems and networks ban them. > In an experiment conducted in 1996 on this system, we allowed users to compile and run their bots. The result was hundreds of megs of disk space became occupied because each user insisted on having their own version of eggdrop uncompressed and untarred in their home directory. All physical memory was in use as ~45 eggdrop processes were running concurrently. The system was basically USELESS and it took 1.5 hours to login if you were patient enough (even from the system console). > The ARPA members called a vote on the issue and the result was almost a resounding unanimous NO. > However, there are times when running a bot is useful, for instance keeping a channel open, providing information or just logging a channel. Basically the bot policy here is a bit relaxed for MetaARPA members. Common sense is the rule. As long as you aren't running a harmful process, such as a hijack bot, warez bot or connecting to a server that does not allow bots, then you may run a bot process. More info about SDF for those who are curious: https://en.wikipedia.org/wiki/SDF_Public_Access_Unix_System https://en.wikipedia.org/wiki/SDF_Public_Access_Unix_System
- bawolff 2y agoAs far as i can tell, this policy is essentially - don't do anything with an llm that would get you banned if you did it manually as a human.