6 ms·
I'm Scared a Stranger Will Call My Novel AI, So I Built GitHub for Words
- 3dedb728-3f77 2mo agoIf the prose is not the same as LLM, then all is good. It is not that it is bad prose, it is that we read it everywhere. We got LLM prose fatigue.
- TimPC 2mo agoLLMs still write middling quality prose. Aspects of it are good but they write in a very tropey manufactured style that isn't something we are fatigued of, it's something that's inherently flawed and disliked.
- eth0up 2mo agoPlease forgive me. I am NOT asserting anything, but I read this (your website), and had it not contained express claims of not being AI, I would have bet a substantial stake that it was AI. Perhaps because AI imitates good writing styles, I am not sure, but it matches many of the signals that for me, indicate AI Edit: The above said, I am seeing in my own writing, artifacts from my interactions with AI. I am proof that, in some cases, AI can be contagious.
- pmg101 2mo agoI agree. My guess is the author either 1. Deliberately used LLM tics in their prose, for artistic and provocative reasons, or 2. Used AI It's quite interesting to reflect on.
- arjie 2mo agoHaha, I have to agree. Their style is either intentionally similar for humorous reasons or they're the kind of writer who was copied. > The words are the crime scene, not the alibi. If the latter, it must suck haha. It reminds me of the fact that I used to wear masks when I had a cold so that I could do things with my friends like help them move or whatever without risking them getting sick. But then one day, a pandemic arrived, and wearing a mask was a political decision. Now if I do, I'm making a political statement unintentionally! Tragic, but it can't be helped. Regardless, I see why they feel the need to make this service. Their post gets hard-flagged by Pangram haha.
- claiir 2mo agoThat’s because it is. https://github.com/dylanreed/dylan.blog/blob/main/docs/BLOG-WRITING-GUIDE.md https://github.com/dylanreed/dylan.blog/blob/main/docs/BLOG-...
- eth0up 2mo agoWell that explains a lot. It sure as hell seems to be. Thanks for the info.
- happytoexplain 2mo agoI've always felt that many people write like LLMs, even before LLMs. It had to come from somewhere. A lot of LLM writing is truly bad - but a lot of it is also just interpreted as an LLM style because of repetition, not because it is inherently bad.
- WCSTombs 2mo ago> That map makes it stupidly easy to tell a human from a machine. Someone who commits 50,000 lines in an afternoon did not type those by hand. Someone doing a couple hundred? Probably real. Simple. Uh, this isn't true, though? I squash commits all the time locally before I publish them. There can be many reasons for that, but at the end of the day, I'm not obligated to share my messy revision history with the world. In the most extreme case, the first commit on one of my projects spans multiple years of development because it was forked from another project that's now pretty irrelevant to it. On the other hand, I've received pull requests that were a couple dozen lines and were definitely LLM-authored. As sort of a counterpoint to the author's goals, IMO the community should extend the benefit of the doubt to writers and artists who explicitly say they don't use AI, because it's just really hard to prove that you're not using AI. Trying to force any kind of system will inevitably affect the quality of what you make, and the art community should not want that.
- bradyd 2mo agoIt's also possible to break the 50,000 lines up into multiple commits and fake the commit dates, so this does nothing to prove a human wrote them.
- Krutonium 2mo agoYea there's literally software that lets you put art into your commit history blog graph on github by making a repo with potentially tens of thousands of dummy commits.
- KolibriFly 2mo agoSquashing commits pretty much turns the whole "green carpet" into meaningless noise
- georgeecollins 2mo agoEntrepreneurs of hacker news: Media companies are experimenting with tools that try to detect if a work contains AI. I don't have the domain knowledge to know if that will work, but it seems like the alternative is to have a system to track the provenance of a piece of art or a written work. If you could say that every change was done by party/ person xy then a company buying or licensing it could comfortably depend xy's assurance that it was original work.
- nh23423fefe 2mo agoConstructing validation claims for useless properties, doesn't render the object with the property valuable?
- retrac 2mo agoGenerative AI could synthesize a plausible commit history, too.
- nonameiguess 2mo agoThis is fundamentally impossible without multiple third parties personally witnessing the act of creation. Consider John Milton was blind and composed Paradise Lost by dictating it to another person. Even if you can prove a human, even a specific human, physically performed the act of typing or writing, that doesn't prove they came up with the words themselves. LLMs can whisper in your ear. Though I don't personally believe this, consider the most widely-read book in the history of western civilization still has contested authorship, with many if not most readers convinced God himself dictated the words to the humans who wrote it down. We have no way to prove or disprove that.
- epiccoleman 2mo agoThe problem with this idea is that it operates on the assumption that consumers will actually care about "provenance" on art. I just don't really buy that the set of people who care about that will be big enough for artists to care about signing up for some kind of provenance-tracking regime. Obviously there will be some people in that category on both sides, but I think it'll be niche. I have a hard time imagining Taylor Swift or Drake or whoever is going to start blockchaining their songwriting so that no one can say they used AI. It won't be a big enough segment of their market to care. (and in fact, the first time Taylor Swift does some big AI-generated video or whatever, all of a sudden a whole bunch of people will suddenly decide they're cool with it)
- nh23423fefe 2mo agoanti-ai people harass each other, and gain nothing because their world model is bad.
- bigyabai 2mo agoIf the AI people had a better world model, then we'd probably know by now. Kinda seems like both "sides" are losing.
- nitwit005 2mo agoUnfortunately, people will truly try to fake the "paper trail". People posted sped up video of them creating art as proof they made it, so people also used AI to try to generate that.
- dandersch 2mo agoDo you have examples? I don't think AI can create convincing timelapses of digital art being made, yet. I know there is nothing stopping it from getting there, but I think artists are happy for any kind of proof-of-work they can produce right now.
- nitwit005 2mo agoThe most casual Google searches show results: https://www.instagram.com/reel/DOvcEN1kXp0/ https://www.instagram.com/reel/DOvcEN1kXp0/ Including a research paper from a few years back: https://inversepainting.github.io/ https://inversepainting.github.io/ Edit: Actually hadn't thought of fake building construction timelapses, and apparently that's also a thing.
- stets 2mo ago> As I’ve mentioned before, we now live in a world where creative people have to prove they created the thing they created. Do you? Just stop caring. I don't think it makes sense to keep asserting that you're an AI vegan in order to please randos. You'll always get hate creating/building anything online. The people that enjoy your work won't care if it's AI or not or they'll believe you.
- happytoexplain 2mo agoEven disregarding the bizarre assertion that a human author could "just stop caring" at will, this is still unrealistic even on a practical level. Enough consumers want to know whether prose is AI that it matters.
- beej71 2mo agoI have yet to be so-accused, but if I ever am, I'm just going to accuse the accuser of being an AI. :)
- scarmig 2mo agoSounds like something an AI would say.
- beej71 2mo agoYou're not wrong, and you're right to push back on that. The speech pattern is real.
- eth0up 2mo agoI think the theme with negative-style statements, eg 'it's not this -- it's this', is authority groveling, ie when you are missing a brain, any successful claim registers toward credibility. It indicates that you know something. I suspect as LLMs evolve, that will disappear. Edit: "you" = the model, not literally you. Poorly written, pardon.
- overgard 2mo ago
- babu_mick 2mo agosuper cool
- deleted 2mo ago[deleted]
- legacynl 2mo agoMy spidey sense is tingling with this post. If someone was really scared that their text would be assumed to be AI, why wouldnt you at least stop using the obvious AI-sign, i.e. the long dashes? Does the meaning of the text really change that much that you MUST use — instead of - ? Also: * No dashes before the 'its not vibe coding post' (or very rare), afterwards lots of dashes. * All pixel art images are AI generated.
- deleted 2mo ago[deleted]
- JoshTriplett 2mo ago> why wouldnt you at least stop using the obvious AI-sign, i.e. the long dashes Em-dashes have a long literary history, which is where AI got them from, and they're genuinely useful in writing. We shouldn't let LLMs ruin that for us in literature. The fact that they've historically been less common in online writing is in part because of deficiencies in input systems, though they're easy enough to type with a compose key or similar, or a typesetting system that interprets `---` as `—`.
- jonesy827 2mo agoRelevant -- https://psychotechnology.substack.com/p/em-dashes-are-fucking-amazing https://psychotechnology.substack.com/p/em-dashes-are-fuckin...
- vunderba 2mo agoI agree but it feels like a bit of a losing battle at this point, and I can't help but be reminded of the history of the Sanskrit swastika.
- legacynl 2mo agoThat might be true. And for pedantic literary critics and writers, it might be very important to use the proper length dash. For anybody else, the difference between the long dash, the longer dash, or the normal dash, is almost invisible and meaningless. > We shouldn't let LLMs ruin that for us in literature. First off all, this is a blog post, not literature. Second of all, the cat is already out of the bag on that one. People already associate it with LLMs, no blogpost is going to change that.
- matteoraso 2mo agoI don't see what this offers that Git doesn't. Even if Git only tracks lines, that's still a good proxy of how many words you wrote. Besides, you can use git diff to see whether an individual character was changed, even if the number of lines stayed the same.
- number6 2mo agoYes, GitHub is a GitHub for words
- KolibriFly 2mo agoNon-techies avoid the CLI like the plague and don't want to mess with .gitattributes for text files. But building a whole web service with a paid subscription just for that is total overkill when free desktop GUIs like Sublime Merge or VS Code exist
- _dwt 2mo agoI'm a little confused in that this post is either (partially?) AI-generated or deliberately written to appear so (including to Pangram). I'm a self-confessed AI hater and yet I agree with the "just write your thing" people on this one. I've watched this play out on Substack and I think there's a mix of a) people genuinely scared of what I believe to be very rare false positives from automatic detection, b) people who use AI (but claim otherwise) trying to use group (a) as cover to argue against automatic detection, and c) the hoi polloi throwing rotting fruit at anyone who uses an em dash. You can safely ignore group (c); they aren't going to be your people even with a fully-attested human changelog witnessed by the Pope and encoded on the blockchain. Group (b) are basically grifters and also to be ignored. Group (a)'s concerns are valid but in the large I don't see a lot of "grey area" writing online; I see a lot of obvious slop. (Possibly the well-done AI assistance goes unnoticed by me, in which case - philosophical concerns aside - I can only think of the old XKCD comic about spam bots - mission accomplished? I don't think so, though.) Back to the OP, I didn't really understand the "meet the enemy" section - did that experiment go anywhere? My results from trying to see if even frontier models can capture voice, style, philosophy from my own technical writing corpus have been terrible to laughable at best.
- stenmorten 2mo agoI recognize myself in that I build instead of publish. This is just fear of rejection, in productivity-clothing. The only way to publish, or to get published, is to put yourself out there. You risk rejection and ridicule. But life is not a spectator sport.
- Keeeeeeeks 2mo agoI think the only way to really prove you wrote something is a camera streaming video and audio from a faraday caged room, and a program like this recording each keystroke Then you can upload that to a private repo after each session. Otherwise people make a program an agent can use to type and retype whole sections like a person
- hyperhello 2mo agoThat would be a major red flag to me — a pointless, highly compressible stream of images that prove nothing in actuality but are easy to define. If your work is being confused with AI, think about why. I never confused Trevanian and Martin Amis and Roald Dahl with AI.
- alainrk 2mo agoInterestingly hackernews commenters are the main reason of this fear
- scarmig 2mo agoThe issue isn't just that cheaters are going to cheat. It's that the project will provide a positive and mostly trustworthy provenance exactly during the period of time when no one cares about or trusts it. However, as knowledge of it diffuses and people care about and trust it more, it will attract more and more cheaters. This negative feedback loop puts an effective cap on how much value people can get out of it.
- gverrilla 2mo agoOnly a small minority is worried about that. And the worries will soften with time, I believe (much money involved, and much power on imaginations). People looking for non-AI text will know where to find it, and I don't think it will involve any kind of 'non-ai receipt' at all.
- saaaaaam 2mo agoI don’t really understand the point of this. Scrivener has a revision history. Even google docs have versioning. And if you’re writing seriously you’ll know that your first draft often bears no resemblance whatsoever to your final draft. Different character names, different structure, different style, different tone. You may well have a whole bunch of things that seem completely different and a whole load of fragments that are written revised, scrapped, unscrapped, re-revised, stitched together, pulled apart and then finally assembled into a working draft. I’m not sure how soemthing that basically counts your words and (maybe) proves you’re not using AI helps with a typical sprawling novel plan.
- s1mon 2mo agoWhy not just use Google Docs? The history is recorded. There are even plug ins which will play it back so you can see it.
- joshuablais 2mo agoI've been using git for words for some time now. Suprised more writers don't!
- phil21 2mo agoSlop is slop. I don't care if it's human generated or AI generated slop. It's all the same to me. All AI did was accelerate the generation of slop so that bit is interesting to discuss and ponder. It's simultaneously uninteresting to complain about a given piece of work. Either it stands on its own merits or does not. If something sucks, I don't care if it sucked because the human did a poor job at it or they lazily used AI. The output is the same to me and all that really matters. I think what really is being exposed here is that so much human generated content before AI use became widespread was slop in of itself. We were just supposed to somehow be fine with it because "effort" might have gone into creating said slop. Nope. Talent matters. The same goes for all the folks complaining about AI generated concert/event posters around where I live. While I can typically discern an AI generated one vs. a human made one - the fact is they all are within a very close quality range to me. This was already low-value slop work most "designers" simply slapped together copying the same basic template. Perfect use for AI imo, since it simply replaced already low-effort work no one actually cared about until they were told to with something far more efficient. It lets the humans who commissioned such pieces do actual work on concert promotion and such vs. multiple hours/days of back and forth with a graphic designer for roughly the same quality output. Now that people are paying attention, the high effort artistic pieces in the category stand out immediately. TBD if anyone actually cares and if this shows up in better event attendance. The final output is what matters. I do not care one lick over how it was achieved, what creative process someone follows, or what tools they use. I don't care that some movie director has an artisanal old-school editing process where they glue physical film together for their cuts. I care about how good the movie is, and if that is their process - awesome! I'm happy for them! If some other director uses a digital means to do the same and cuts down their editing time by 80% I really do not care and I'm happy for them too!
- spookymutation 2mo ago> The part that makes me want to chew glass is that none of this would be a problem if people were just honest about what they used to make a thing. But that’s apparently too much to ask, because cheaters gonna cheat. And yet the author does not seem to be forthright about their use of LLMs to write their "author website", blog posts, Vellum"P"roof website, and presumably Vellum"p"roof (capitalization stylized differently on the website and blog) itself. > Your prose is never fed to an AI, never used to train a model, and never handed to anyone else. And yet their is a distinct lack of a TOS/etc. codifying this promise. The site does accept payments via Stripe. This almost seems satirical.
- claiir 2mo agoSaw quite a few a Claude-isms, like: “It’s git commits, minus..” The blog is AI-created—the commits all have “Claude” as a co-author. [1] The “BLOG-WRITING-GUIDE.md” used to write this post gives it away. [2] And the product the blog is advertising (“VellumProof”)—that website is 100% Claude too. [3] Perhaps this is some social experiment? Don’t really get it. [1]: https://github.com/dylanreed/dylan.blog/commit/a7327d0968692a6ad48e08d4b69cbb8da2b7db95 https://github.com/dylanreed/dylan.blog/commit/a7327d0968692... [2]: https://github.com/dylanreed/dylan.blog/blob/main/docs/BLOG-WRITING-GUIDE.md https://github.com/dylanreed/dylan.blog/blob/main/docs/BLOG-... [3]: https://vellumproof.com https://vellumproof.com
- KolibriFly 2mo agoIt's pretty ironic to pitch protection against AI accusations with a post that was completely generated using Claude prompts
- VCFundedGenYer 2mo agoThe irony that this is AI written.....
- pickleglitch 2mo agoI don't have to worry about this because almost no one ever reads anything I write. Even if they did, I don't give a shit.
- TimPC 2mo agoIt seems to me that recent changes have dramatically exacerbated this problem. Pangram is still claiming their extremely low false positive rate from before adding humanization detection and now very frequently detects actual human works as humanized AI.