25 ms·
I am deeply enjoying this comment thread - it's a bit of a Barium Meal [0] for determining how many people read (a) the headline, (b) the first paragraph, or (c
by meredydd 6y ago
I am deeply enjoying this comment thread - it's a bit of a Barium Meal [0] for determining how many people read (a) the headline, (b) the first paragraph, or (c) the whole thing before jumping straight into the compose box.
Having read to the bottom, the quality of text generation there absolutely blew me away. GPT-2 texts have a somewhat disconnected quality - "it only makes sense if you're not really paying attention" - that this article lacks entirely. Adjacent sentences and even paragraphs are plausible neighbours. Even on re-reading more closely, it doesn't feel like the world's best writing, but I don't notice major loss of coherence until the last couple of paragraphs. I am now really curious about the other 9 attempts that were thrown away. Are they always this good?!
[0] https://en.wikipedia.org/wiki/Canary_trap#Barium_meal_test https://en.wikipedia.org/wiki/Canary_trap#Barium_meal_test
- jcahill 6y agoGPT-3 is a neat party trick. But the things that'll be done with web archives* in the next 20y will make it look like the PDP-8. ~love, a web archivist * GPT-3 is trained on one
- deleted 6y ago[deleted]
- h0p3 6y agoNo pressure: feel free to ignore me, please. Would you mind elaborating? I'm interested in what you have to say (and, of course, feel free to say it privately if you prefer). I would like to even hear your dreams, wild speculations, or gut feelings about the matter.
- jcahill 6y agoSure, what do you want to know? I currently work on synbio × web archival. Some of us are cooking up futuretech aimed at storing all of IA (archive.org) in a shoebox. Others are working on putting archival tools in more normal web users' hands, and making those tools do things that people tend to value more in the short-term, like help them understand what they're researching, rather than merely stash pages. My ambitions for web archives are outsized compared to other archivists, but I'm fine with that. I'm looking beyond web archives as we currently understand them toward web archives as something else that doesn't quite exist yet: everyday artefacts, colocated and integrated with other web technology to an extent that they serve in essential sensemaking, workflow, and maybe security roles. Right now, some obvious, pressing priorities are (a) preserving vastly more content and (b) doing more with the archives themselves. A: The overwhelming majority of born-digital content is lost within a far narrower time-slice than would admit preservation at current rates, and data growth is accelerating beyond the reach of conventional storage media. So, for me, the world's current largest x is never the true object of my desire. I'm after a way to hold the world that is and the world to come. Ideally, that world to come is one where lifelong data stewardship of everything from your own genome to your digital footprint is ubiquitously available and loss of information has been largely rendered optional. This, of course, requires magic storage density that simply defies fundamental limitations of conventional storage media. I'm strongly confident that we're getting early glimpses of the first real Magic contenders. All lie outside, or on the far periphery of, the evolutionary tree that got us the storage media we have today. For instance, I'm running an art exhibition that involves encoding all the works on DNA. B: Distributed archival that comes almost as naturally as browsing is well within reach, and with that comes some very new potential for distributed computation on archives. One hand washes the other. One important thing to realize here is that, in many cases, you can name a very small handful of individuals as the reason why current archival resources exist. GPT-3 is cracking the surface by training on data produced by one guy named Sebastian, for instance. …i'm sorta tired and have to respond to something about every twitter snapshot since June being broken, though, so I'll pick this back up later.
- greyface- 6y agoThis is an interesting thought. GPT-3 used 45TB of raw CommonCrawl data (which was filtered down to 570GB prior to training). The Internet Archive has 48PB of raw data.
- GreenHeuristics 6y agoThat 48PB is mostly just old video game roms and isos though
- ypcx 6y agoThe transformer model as presented in GPT-3 may be a few tweaks away from a human-acceptable reasoning, at which point we may realize that human brain is just a neat party trick as well. This may come difficult for some people to internalize, especially those who understand the technology in depth. Because it means that the medium of our reality is the consciousness.
- yomly 6y agoI don't fully understand what you're getting at here... Basically the brain and "consciousness" isn't as fancy as we think?
- walleeee 6y agoWas this comment generated by GPT-3?
- Naracion 6y agoI doubted that as well, but I don't think it is--at least it's not a simple copy paste. There's an emphasis on _is_ in the last sentence which I don't think the algorithm could have generated. However that makes one wonder if it can also learn to generate emphases, and if so, how would it format? With voice generation it can simply change its tonality but with text generation it has to demarcate it in some way--does the human say "format the output for html", for instance?
- visarga 6y ago> Because it means that the medium of our reality is the consciousness. I agree. The environment - as the source of learning and forming concepts, is the key ingredient of consciousness, not the brain.
- lucidrains 6y agoExactly.
- sytelus 6y agoYou are confusing pattern matching with reasoning. If your brain was replaced by GPT-3 model and you were cast away on a distant island, I highly doubt you will be able to perceive, plan and prosper during your survival against all the calamity nature would through at you.
- scoot_718 6y agoHopefully in a way that secures some funding for those making archives of the web.
- jcahill 6y agoI'm running the Coronavirus Archive. Largest thematic archive on the pandemic, since January. I'm also teaching community biolab techniques to people in parts of the world without ready access to commercial COVID-19 test kits, on all but zero resources at this point. I could use… what's the word? I think it's more funding.
- zamadatix 6y agoOr maybe we're all bots too and you're the only real HN user! I agree, responses are almost as interesting as GPT-3. And this place has always felt like one of the better when it comes to people reading past the titles!
- JBiserkov 6y ago"Every account on reddit is a bot except you." https://www.reddit.com/r/AskReddit/comments/348vlx/what_bot_accounts_on_reddit_should_people_know/ https://www.reddit.com/r/AskReddit/comments/348vlx/what_bot_...
- Alex3917 6y ago> Even on re-reading more closely, it doesn't feel like the world's best writing, but I don't notice major loss of coherence until the last couple of paragraphs. I guessed it was fake before getting to the end, not from the content, but from the fact that all the sentences are roughly the same length and follow the same basic grammatical patterns. Real people purposely mix up their sentence structure in order to keep the readers engaged, whereas this wasn't doing that at all. Still very impressive though; if not for the fact that the post was about computer generated content I probably wouldn't have noticed.
- dannyw 6y agoIf that’s the only thing that separates this from human writing; I’m sure it can be influenced easily.
- Alex3917 6y ago> I'm sure it can be influenced easily. Maybe. Right now this reads like a glorified shopping list. It's coherent, but actually sounding human also requires a theory of mind. E.g. I explain here why it's possible for written statements to be objectively insightful, informative, interesting, or funny, but objectively in a way that's relational to other information or beliefs. The implication being that statements are only going to seem subjectively funny or insightful (or whatever) to others who have that knowledge or those beliefs, which means that you can't reliably create those subjective experiences in a reader without having some sort of theory of mind for them. I guess you can create content that's funny or insightful relative to that content itself, but that's not especially useful. It's entertaining at the time, but the experience is more like seeing a movie that you laugh a lot during but then leave and are kind of like what was the point? It's an empty experience because it wasn't transformative. I definitely don't think it's impossible, but I also don't think it's a matter of just adding a couple more if-else statements. https://alexkrupp.typepad.com/sensemaking/2010/06/how-writing-creates-value-.html https://alexkrupp.typepad.com/sensemaking/2010/06/how-writin...
- roenxi 6y ago
- m3kw9 6y agoIs exactly the issue. You still need humans to check it before releasing an output. It can only be what the author says “Bitcoin” level of implication if it can get things probably needing at least “99%” quality and correctness.
- joe_the_user 6y agoOK, I read the first sentence and it sounded like a typical poor-quality marketing-blather article so I came here and read your post. I then reread it and it indeed read like a weird, rambling, incoherent article. Looking at it closely, it had a good many contradictory, meaningless and incoherent sentences.("It is a popular forum with many types of posts and posters.") The headline, however, seemed about right. It's true the nonsense in this article is a bit different than the nonsense of a GPT-2 article. But the thing GPT-2 paragraphs sound pretty coherent 'till they suddenly go off the rail. This is more like an article that was never quite on the rails and so it's slightly more internally cohesive. But not "better". Maybe the article just reflects the author's style. Anyone have a GPT-3 test site link?
- m3kw9 6y agoProblem was I lost interest half way because it lost my interest after the 2nd paragraph. For those that say it was good till the last few is really pretending to understand what it said. It really did not made much sense.
- wcoenen 6y agoIt's ironic that you're critiquing the algorithm for not making sense, while contradicting yourself in your first sentence. Did you lose interest "half way", or "after the second paragraph"? It can't be both.
- qotgalaxy 6y agoI'm assuming that the comment that you are replying to was generated by GPT-3. Maybe this one too?
- necovek 6y agoWouldn't "after the second paragraph" be "half-way" for a four paragraph piece? :D But you are right, it can't be both in the context of this article :)
- ricksharp 6y agoThose types of contradictions are what made me suspect the article was generated. Now, I’m not so sure :)
- blueblimp 6y agoI'm finding with secret GPT-3 output that I often find it boring before realizing it's GPT-3. I might even be getting to recognize its wordy, dull, cliche-ridden, borderline-nonsensical style. It's remarkably good at passing as human writing of no value whatsoever.
- visarga 6y agoThe fast jump in quality from GPT2 to 3 is more important than the current level of GPT3. Maybe next year it will be not-boring.
- yyyk 6y agoI'm not a very good test case. I briefly skimmed (not expecting very much from a Bitcoin-themed article), read the end, and only then read more carefully. So my first read was brief and biased, and my second was very biased. That said, the content of the computer-generated parts doesn't make much sense even for a Bitcoin-influenced article (what would be the point of paraphrasing your previous post in a forum on a regular basis, and how does this not get one very quickly banned?), but the grammar is far far better than previous attempts - it reads like Simple English wiki.
- woko 6y ago> I briefly skimmed (not expecting very much from a Bitcoin-themed article), read the end, and only then read more carefully. It sounds to me like you must be an academic, or someone with good habits for being efficient at reading articles.
- Reimersholme 6y agoI do think the starting paragraphs with background information and contextualization is really impressive. I can definitely see for example Wikipedia injecting a knowledge graph and having a system like this output full articles in X languages automatically in the near future.
- sidcool 6y agoThis is so true. I had a high regard for HN comments till quite recently.
- sangnoir 6y agoI take it you've never had your comment - which happens to be on a subject of your expertise or lived experienced - downvoted because simply because it doesn't feel "truthy" enough for HNs demography. I have long since adopted a healthy disregard of HN comments in other areas that I'm no expert in. I still haven't found a way to monetize that though; that is my holy grail.
- thewarrior 6y agoI've started working on a version of GPT-2 which generates English text. The purpose of this is to improve its ability to predict the next character in a text, by having it learn 'grammatical rules' for English. It already works well for predicting the next character when it has seen only a small amount of text, but becomes less accurate as the amount of training text increases. I have managed to improve this by having it generate text. That is, it creates an 'original' piece of text about 'topic x', then a slightly altered version of this text where one sentence has a single word changed, and this process is repeated many times (about a million). It seems to quickly learn how to vary sentences in a way that seems natural and realistic. I think the reason this works is because it reduces the chance that the grammar it has learned for one specific topic (e.g. snow) will accidentally be transferred to another topic (e.g. dogs). Of course, this all means nothing unless it actually learns something from the process of generating text. I haven't tried this yet, but the plan is to have it generate text about a topic, then have a second GPT-2 system try to guess what that topic is. If the resulting system is noticeably better at this task, then we know the process has increased its ability to generalize. One potential issue with this approach is that the text it generates is 'nonsensical', in that it is almost like a word-salad. Although this is a standard problem with neural nets (and other machine learning algorithms), in this case the text actually is a word-salad. It seems that it has learned the rules of grammar, but not the meaning of words. It is able to string words together in a way that sounds right, but the words don't actually mean anything. Plot twist: This comment was generated by GPT-3 prompted with some of the comments in this thread.
- cs702 6y agoThe thing that kills me is that to the vast majority of human beings the nonsensical technobabble above is probably indistinguishable from real, honest, logically consistent technobabble.[a] Soon enough, someone will replicate the Sokal hoax[b] with GPT-3 or another state-of-the-art language-generation model. It's not hard to imagine GPT-3 writing a fake paper that gets published in certain academic journals in the social sciences. [a] https://en.wikipedia.org/wiki/Technobabble https://en.wikipedia.org/wiki/Technobabble [b] https://en.wikipedia.org/wiki/Sokal_affair https://en.wikipedia.org/wiki/Sokal_affair -- here's a copy of Sokal's hoax paper, "Transgressing the Boundaries: Towards a Transformative Hermeneutics of Quantum Gravity:" https://physics.nyu.edu/faculty/sokal/transgress_v2/transgress_v2_singlefile.html https://physics.nyu.edu/faculty/sokal/transgress_v2/transgre...
- c3534l 6y agoI read the beginning, went "what the fuck is this guy on about? Get to the point" and then came to check the comments to if this was anything interesting or worth reading, saw your comment, and skimmed the end bit. Overall I'm pleased with my process as its an efficient way to find out which articles are worth reading. But it was also clear to me that the author had difficulty making a clear point or had a goal in his writing. I skipped through it for a reason and I suspect many other people did as well.
- ALittleLight 6y agoFrom your comment it's not clear to me if you realize the author of the article is GPT-3.
- mdoms 6y agoIt's very clear to me that he does not. But he does an excellent job of making GP's point.
- jjoonathan 6y agoErr, it seems very clear to me that he does realize GPT-3 is the author, and that it was easily caught by his bullshit filter. Which was my experience too -- but I am less dismissive. I regularly see human-produced bullshit get very far with less coherence than these examples from GPT-3. GPT-3 isn't AGI, but it's weapons-grade in a way that GPT-2 wasn't.
- AnimalMuppet 6y agoOK, but a weapons-grade BS generator is not what the world needs right now...
- jjoonathan 6y agoReady or not, here GPT-3 comes.
- 6y ago
- DonHopkins 6y agoIt reminds me of the incoherent demented ramblings we've all been been hearing (but hopefully not following as medical advice) for the past several years.
- OOPMan 6y agoThanks to this comment I actually read the blog post. It was relatively good, although I began to suspect it was GPT3 generated about halfway through (partially because the style felt a bit stiff but also just out of a shayamalan-what-a-twist 6th sense of mine that was tingling)
- dTal 6y agoIt is a bit unfortunate that your comment is now at the top - it spoils the test :) Saying that, I briefly saw the first sentence of your comment and went to read the article with the idea that trickery was afoot, specifically guessing correctly the nature of the article. And yet, even then, on the back foot... it fooled me. Incredible.
- simonebrunozzi 6y agoAh, but you are almost spoiling the end with your second paragraph! :) I agree with you. I suspect few people have read until the end to realize that, in fact, ...
- classified 6y agoEliza could do this better. Or, just use a Markov chain that has read enough corporate PR bullshit. It's just sad how many people use this "AI" meme to fulfill their need to worship something.