3 ms·
This argument seems more like the data generated was bad. There are examples where AI has surpassed humans by using simulated data (AlphaZero - where it played
by patrickhogan1 2y ago
This argument seems more like the data generated was bad. There are examples where AI has surpassed humans by using simulated data (AlphaZero - where it played against itself to become the best at Go).
It also seems to happen most on small networks. Which makes sense.
Additionally, humans create simulated stories like Dune, Lord of the Rings, or Harry Potter, which introduce fictional concepts, yet these stories still result in trainable data.
- thaumasiotes 2y ago> Additionally, humans create simulated stories like Dune, Lord of the Rings, or Harry Potter, which introduce fictional concepts, yet these stories still result in trainable data. No, they don't, not in any sense where they are "simulated data". Dune is simulated data about what life would be like on Arrakis, and if you train a model to make predictions about that question, your model will be worthless trash. (Doesn't matter whether you train it on Dune or not.) Dune is real data about how English is used.
- ninetyninenine 2y agoIt’s also data around science fiction. With broad spectrum data from both dune and contextual data most LLMs know that dune is from a fictional novel.
- raincole 2y ago> humans create simulated stories like Dune, Lord of the Rings, or Harry Potter People really anthropomorphize LLM to a full circle, don't they?
- patrickhogan1 2y agoSo you are saying that if I generate stories on different worlds as new data, a model cannot learn from that? This isn't anthropomorphizing - it's generating data. Generating data is not a uniquely human endevor. What created Mars? That is data. What created star systems we cannot see?
- banku_brougham 2y agothis is not a serious argument, please forgive me for saying
- dudeinjapan 2y agoThank you for making this comment, because it exposes some logical gaps. Firstly, Go, Chess, and other games have objective rules and win criteria. (There is no “subjective opinion” as to whether Fischer or Spassky won their match.) Language, the output of LLMs, does not have an objective function. Consider the following to sentences: “My lips, two blushing pilgrims, ready stand.” “My red lips are ready to kiss your red lips.” Both are grammatically correct English sentences and both mean basically the same thing, but clearly the former by Shakespeare has a subjective poetic quality which the latter lacks. Even if we make evaluation rules to target (for example, “do not repeat phrases”, “use descriptive adjectives”, etc.) AI still seems to favor certain (for example “delve”) that are valid but not commonly used in human-originated English. There is then a positive feedback loop where these preferences are used to further train models, hence the next generation of models have no way of knowing whether the now frequent usage of “delve” is a human-originated or AI-originated phenomenon. Lastly, regarding works of fiction, the concern is less about the content of stories—though that is also a concern—but more about the quality of language. (Consider above alternate take on Romeo and Juliet, for example.)
- patrickhogan1 2y agoSo you are arguing that the world does not have objective rule criteria, like Physics? And that an AI could not model the world and then run simulations and have each simulation generate data and learn from that, similar to AlphaZero. Here is a possible objective win environment Model complex multicellular organisms that become capable of passing the Turing test that also self-replicate.
- dudeinjapan 2y agoMy argument here is narrowly scoped to human language and literature. (We already know the objective rule criteria of life is 42.) It may very well be possible for an AI to read all of literature, and figure out what makes Hemingway, Tolstoy, and Dylan "good writing" vs. "bad writing". That has not yet been achieved. The problem, as the OP implies, is that by polluting the universe of literature with current-gen AI output, we may making the task of generating "good writing" in the future harder. Then again, maybe not. Perhaps we have enough pre-AI works that we can train on them versus the mountains of AI generating schlock, and determine the objective function.