11 ms·
AlphaWrite: AI that improves at writing by evolving its own stories
- passwordoops 1y agoAppropriate that the original title misspells "Writing": >AlphaWrite: Inference time compute Scaling for Writting
- SamBam 1y agoI found the entire first sentence nearly unreadable: "Large languagenference time compute Scaling for Writing models have demonstrated remarkable improvements in performance through increased inference-time compute on quantitative reasoning tasks, particularly in mathematics and coding" Am I just out of the loop on the current jargon, or is that indeed a terribly-written first sentence?
- fside 1y agoThe workflow here feels pretty natural, just using the AI to help with the boring parts and speed things up. I like the idea of treating it as a tool, not a replacement.
- xnx 1y agoNote: Not associated with Google Deepmind (AlphaFold, AlphaGo, AlphaEvolve, etc.)
- kotaKat 1y agoNor associated with AlphaSmart word processors or their AlphaWrite application. Sigh.
- Applejinx 1y agoIn this genre do you really expect a lot of concern for intellectual property or the ability to identify the source of anything?
- OhioMan2943 1y agoThough they obviously want to be to the point of infringing. It's the modern AI legal bubble so they'll never have to deal with the legal consequences.
- zorked 1y agoIf there is something that I would like AI to never touch, it's that. Please stop making the world worse.
- nougati 1y agoLike it or not, this stuff will happen. Might as well have curiosity about it
- contagiousflow 1y agoTechnological change doesn't happen independent of culture. Stop with the technological determinism.
- smitty1e 1y agoThe proper counter will be cultural determinism, if sufficient people insist upon supporting human writers crafting real books.
- vidarh 1y agoI'd argue that technological change that is a simple enough iteration of existing technology that it can be carried out by small groups of people will happen independent of culture with a very high probability once a population is large enough. In this case, so many people are curious as to whether we can make this work and/or see financial implications that this will happen irrespective of whether wider culture rejects it.
- contagiousflow 1y agoPrioritizing profit above all is still a culture though. And it is the choice of a people to give a small group with large amount of money huge cultural influence.
- deadbabe 1y agoYou cannot stop people from making the world worse or better. The best you can do is focus on your own life. In time many will say we are lucky to live in a world with so much content, where anything you want to see or read can be spun up in an instant, without labor. And though most will no longer make a living doing some of these content creation activities by hand and brain, you can still rejoice knowing that those who do it anyway are doing it purely for their love of the art, not for any kind of money. A human who writes or produces art for monetary reasons is only just as bad as AI.
- suddenlybananas 1y agoI can't imagine LLMs are good judges of good writing.
- itishappy 1y agoThat's the core of my concern too. Be interested to see what happens if you feed the ranking algorithm a list of the most popular books and a list of the most impactful books. Something tells me this will be a lot more interested in Chuck Tingle than Kafka.
- CuriouslyC 1y agoLLMs are fairly good judges of writing, in fact they're better at evaluating writing than they are at actually writing. I use Gemini as a beta reader, and I've had a lot human beta readers look at the same material, and Gemini consistently gives significantly better than average feedback, though it's stronger at structural and prose evaluation and weaker at emotional and "wishlist" style feedback as you would probably expect.
- riskable 1y agoHow are you using it as a beta reader? What prompts do you use? I'd love to try it.
- CuriouslyC 1y agoJust dump your manuscript into google's aistudio, and tell it you'd like it to serve as a beta reader/editor, and tell it what your objectives are with your manuscript so it can give you targeted feedback.
- riskable 1y agoI've pasted whole chapters (of my own writing) into ChatGPT and Claude that I know need drastic improvements. Basically, they were first draft, "get the concept down; don't think too hard" paragraphs with occasional run-on sentences and whatnot. This is my very first novel (ever) so of course the initial draft is going to be bad. Both ChatGPT and Claud always say something like, "a few grammar corrections are needed but this is excellent!" So yeah: They're not very good at judging the quality of writing. Even with the "we're trying not to be sycophants anymore" improvements they're still sycophants. For reference, I mostly use these tools to check my grammar. That's something they're actually quite good at. It wasn't until the first draft was done that I decided to try them out for "whole work evaluation".
- saberience 1y agoWhy would we ever want to build something like this, unless your goal is to have fiction writers make even less money than they already do. Just stop, please. Try and automate some horrible and repetitive drudgery. Do you want to live in a world where humans no longer do any creative work? It’s grotesque.
- frozenseven 1y agoI want to live in a world with more options and freedom to choose. If somebody wants to build and explore with these AI systems, you can't stop them.
- Gerardo1 1y ago> I want to live in a world with more options and freedom to choose And you think that AI will be beneficial to that want?
- frozenseven 1y agoYes, undoubtedly.
- soulofmischief 1y agoFalse dichotomy, overextended outrage and straw man fallacies all rolled into one. Impressive.
- itishappy 1y agoThose examples seem quite unrelated to one another. The first reads as admitting intentional fraud and deceit, the second reads like dealing with imposter syndrome. I'd love to know the prompt. Also, not sure how you can judge a style to be clearly better than another. The workflow of generating a bunch of stories in the style of different authors and then voting on a favorite just seems like picking a favorite author. Will the system ever prefer short, hard-hitting sentences? Sure enough, convergence is a noted behavior.
- notahacker 1y ago> Those examples seem quite unrelated to one another. The first reads as admitting intentional fraud and deceit, the second reads like dealing with imposter syndrome. I'd love to know the prompt. Yeah. And to read the rest of each of the stories it generated... Both paragraphs are simply short excerpts which involve no actual narrative, never mind the stuff that LLMs are typically weak at (maintaining consistency, intricate plotting and pacing, subtlety in world and character building) which in the context of stories are far more important to improve than its phrasing. The fact that the "improvement" apparently eliminates a flaw in the first passage ("gentle vibrations that vibrated through my very being" is pretty clunky description unlikely to be written by a native human; both paragraphs are otherwise passable and equally mediocre writing) by implying apparently completely different (and frankly less interesting) character motivations makes me doubt that it's actually iteratively improving stories rather than just spitting out significant rewrites which incidentally eliminate glaring prose issues.
- tamassimond 1y agoYeah as we mention in the blog it's really hard to eval on short passages. If you go on the Github can see longer stories where the change is more noticeable. Both those stories are from the same prompt
- riskable 1y ago> Also, not sure how you can judge a style to be clearly better than another. This one is actually easy: The writing style used for a horror is different than what you'd use for a romance novel. Example: If you give it a prompt that asks the AI to generate something in the style of a romance author but the rest of the prompt is describing a horror or sci-fi story you'll end up with something that most people would objectively decide, "ain't right."
- mynti 1y agoseems like this is just reward hacking the llm as a judge. this does not give you a story humans will be more likely to read imho
- mock-possum 1y agoThis feels like a lot of fluff, without some solid examples of the results - the one example of generated prose that they do provide is pretty unimpressive, it reads like… well, like an LLM wrote it.
- SamBam 1y agoIndeed. Also the whole thing is just "apply an Evolutionary Algorithm to stories." The only interesting question is whether the LLM that decides the story's fitness rating (which they call Elo, despite seeming to have nothing to do with the actual Elo ranking system) can mimic a human's rating. Given the brief example, it's not clear that it can, since it seems no better.
- tamassimond 1y agoTo clarify we use an Elo ranking system to update models scores, so if you loose to a higher rated story you don't loose as much Elo ranking. Definitely agree with LLM judge criticism though it's still an open questions of how we can make them better. Using the repeated story comparison judging system does help make them more consistent. A good rubric helps make them more human like as-well. The really big question is how large is the generator verifier gap between creating stories and marking them
- gabriel666smith 1y agoI'm struggling a bit to understand the difference between the reported results in the blog post and the examples in the Github. The blog states: > "Alpha Writing demonstrates substantial improvements in story quality when evaluated through pairwise human preferences. Testing with Llama 3.1 8B revealed: 72% preference rate over initial story generations (95 % CI 63 % – 79 %) 62% preference rate over sequential-prompting baseline (95 % CI 53 % – 70 %) These results indicate that the evolutionary approach significantly outperforms both single-shot generation and traditional inference-time scaling methods for creative writing tasks." But in all of the examples using Llama 3.1 8B on the Github that I could find, the stories with the top 5 highest final 'ELO' are all marked elsewhere as: "generation_attempt": null Where the 'variant' stories, which I take to be 'evolved' stories, are marked: "generation_type": "variant", "parent_story_id": "897ccd25-4776-4077-a9e6-0da34abb32a4" IE - none of the 'winning stories' have a parent story; they seem to have explicitly been the model's initial attempt. The examples seem to prove the opposite of the statement in the blog post. Perhaps 'variants' are slightly outperforming initial stories on average (I don't have time to actually analyse the output data in the repo), though it seems unlikely based on how I've read it (I could be wrong!) and this might be borne out with far more iterations. However, a really important part of creative writing as a task is that you (unfortunately) only get to tell a story once. The losing variants won't ultimately matter. So, if I've read it correctly, and all the winning stories are 'not evolved' - from the initial prompt - this is quite problematically different from the blog's claim that: > "we demonstrate that creative output quality can be systematically improved through increased compute allocation" Super interesting work - I'd love to be told that I'm reading this wrong! I was digging through in such detail to actually compare differently-performing stories line-for-line (which would also be nice to see - in the blog post, perhaps).
- tamassimond 1y agoJust to clarify slight misunderstanding the variants without parent ID aren't from the initial batch it just didn't carry over to the next batch. You can see "897ccd25-4776-4077-a9e6-0da34abb32a4" emerges from batch 5. Apologies probably should make this clearer. Appreciate feedback on blog post!
- seaourfreed 1y agoThis open source project will grow a ton, if it had a license on it. Can you add an MIT license?
- tamassimond 1y agoThanks for feedback added MIT license
- dgeiser13 1y agoNo. There's no thinking behind AI. If anything it throws shit against the wall repeatedly until a human steps in and says "that seems to be an improvement".
- SkiFire13 1y agoThis completely misses the point of reinforced learning. The reward condition needs to be representative of what you want (e.g. in chess that would be winning). Using a LLM as a judge means you will ultimately optimize for stories that are liked by the LLM, not necessarily for stories that are liked by people. For this to work the other LLM needs to be as close to a human as possible, but this is what you were trying to do in the first place!
- proof_by_vibes 1y agoAs a playwright, I've certainly thought about AI impacting the art. In fact, it was the very eloquence of chatgpt's output that initiated all of this mania in the first place: not only was chatgpt able to explain to me gauge theory with surprising accuracy, it was able to do so using perfect Elizabethan english—exactly as I had instructed it to. There is a missing ingredient that LLMs lack, however. They lack insight. Writing is made engaging by the promise of insight teased in its setups, the depths that are dug through its payoffs, and the revelations found in its conclusion. It requires solving an abstract sudoku puzzle where each sentence builds on something prior and, critically, advances an agenda toward an emotional conclusion. This is the rhetoric inherent to all storytelling, but just as in a good political speech or debate, everything hinges on the quality of the central thesis—the key insight that LLMs do not come equipped to provide on their own. This is hard. Insight is hard. And an AI supporter would gladly tell you "yes! this is where prompting becomes art!" And perhaps there is merit to this, or at least there is merit insofar as Sam Altman's dreams of AI producing novel insights remain unfulfilled. This condition notwithstanding, what merit exactly do these supporters have? Has prompting become an art the same way that it has become engineering? It would seem AlphaWrite would like to say so. But let's look at this rubric and evaluate for ourselves what else AlphaWrite would like to say: ```python # Fallback to a basic rubric if file not found return """Creative writing evaluation should consider: 1. Creativity and Originality (25%) - Unique ideas, fresh perspectives, innovative storytelling 2. Writing Quality (25%) - Grammar, style, flow, vocabulary, sentence structure 3. Engagement (20%) - How compelling and interesting the piece is to read 4. Character Development (15%) - Believable, well-developed characters with clear motivations 5. Plot Structure (15%) - Logical progression, pacing, resolution of conflicts""" ``` It's certainly just a default, and I mean no bad faith in using this for rhetorical effect, but this default also acts as a template, and it happens to be informative to my point. Insight, genuine insight, is hard because it is contingent on one's audience and one's shared experiences with them. It isn't enough to check boxes. Might I ask what makes for a better story: a narrative about a well developed princess who provides fresh perspectives on antiquated themes, or a narrative about a well developed stock broker who provides fresh perspectives on contemporary themes? The output fails to find its audience no matter what your rubric is. And here lies the dilemma regarding the idea that prompts are an art: they are not. The prompts are not art by the simple fact that nobody will read them. What is read is what all that is communicated and any discerning audience will be alienated by anything generated by something as ambiguous as a English teacher's grading rubric. I write because I want to communicate my insights to an audience who I believe would be influenced by them. I may be early in my career, but this is why I do it. The degree of influence I shall have measures the degree of "art" I shall attain. Not by whether or not I clear the minimum bar of literacy.
- pre2w 1y ago[dead]