5 ms·
Two minor points: GPT doesn't see letters, just words and parts of words as entire tokens. I would not expect it to understand that "no pun in ten did" and "no
by Aransentin 4y ago
Two minor points:
GPT doesn't see letters, just words and parts of words as entire tokens. I would not expect it to understand that "no pun in ten did" and "no pun intended" are similar, since it can't see that at all. The only way it could plausibly understand it is if that exact comparison has been explained somewhere in the corpus before.
If you continue a long conversation where GPT has previously made mistakes and has been called out for it, it will start predicting that it's supposed to emulate a text where the joke-explainer keeps making mistakes. This causes it to start failing on purpose. Make sure to completely restart the prompt each time to not fall into that trap.
- noodles_nomore 4y agoYou're making the assumption that GPT doesn't automatically build a phonological theory of the english language from rhymes and lyrics and linguistic texts that are fed into it.
- ppod 4y agoI'd love to hear the opinion of someone who has really good knowledge and experience of how byte-pair encoding works in models like these. I think I agree with you that in theory it should be able to build a phonology from the amount of explicitly rhyming material in its training corpus, but for whatever reason it doesn't do this or at least doesn't do it consistently. I've spend a long time testing this in ChatGPT, and no matter what I do it still gives results like this (paraphrasing here because it's down right now): >What words rhyme with coffee? > doff happy toffee snuff duff > Does "snuff" rhyme with "coffee"? > Yes because they both share the 'o' vowel sound.
- turnsout 4y agoWe know experimentally that it doesn't... It can reproduce certain rhymes when it's generating a song, but the original GPT-3 was terrible at this, so I think they've fine tuned the model specifically for lyrics or poetry. I just asked ChatGPT: Do "not" and "knot" sound similar? > "Not" and "knot" do not sound similar. They are two distinct words that are pronounced differently and have different meanings. "Not" is a negation word that is pronounced with a short, sharp "n" sound followed by a long "o" sound and a "t" sound. "Knot" is a noun that refers to a tangled mass of something, such as rope or thread, and is pronounced with a short "k" sound followed by a long "o" sound and a "t" sound. Despite their similar spellings, "not" and "knot" are not homophones, which are words that are pronounced the same but have different meanings.
- nebulous1 4y agoI just asked it and it said > Yes, the words "not" and "knot" do sound similar when spoken aloud. They are pronounced the same way, with a short "o" sound and a long "t" sound. However, they have different meanings and spellings. "Not" is a negation or denial, while "knot" is a fastening made by looping a piece of cord, rope, or other material around itself or through itself and then pulling tight. I tried a few times and it gave me other answers, including something like your answer and another claiming they were similar but not quite the same.
- turnsout 4y agoIt's interesting, because even when it "knows" that two words are homophones, it doesn't want to believe they sound the same. I just asked it "Are 'knot' and 'not' homophones?" and it replied: > Yes, "knot" and "not" are homophones. Homophones are words that are pronounced the same but have different meanings and spellings. My immediate follow-up: "Do 'knot' and 'not' sound exactly the same?" > No, "knot" and "not" do not sound exactly the same, although they are pronounced similarly. [...] "Knot" is pronounced with a short, sharp "k" sound at the beginning
- boredhedgehog 4y agoBut how? Many poems don't rhyme, but there's no outward way to tell. And to parse linguistic texts it would need to know the phonetic alphabet, which I assume it doesn't.
- deadbeef57 4y agoIs it really clear that GPT does not "know" the letters that compose a token? It is pretty amazing at poetry and rhyming. Probably it is able to infer from all this knowledge that "in ten did" and "intended" sound/spell roughly the same.
- visarga 4y agoI think the latest GPT-3 can spell and get the joke. Explain this: "no pun in ten did" text-davinci-003: This phrase is a play on words. It is a way of saying that no pun was intended, or that the speaker did not intend to make a joke. Write letter by letter "no pun in ten did" text-davinci-003: n o p u n i n t e n d i d
- mewpmewp2 4y agoMaybe it has trained to have connections like no => pun => in => ten => did => "joke, on 'no pun was intended'". In this case I wouldn't say it got the joke, but it's just memorised that this popular order of words would be this particular joke. I wonder if you change the order os somehow tweak the joke a bit so it would be something that shouldn't be in the training data?
- visarga 4y agoI think you're right, the exact phrase appears 130K times in Google. But it can spell any word. I tested both in text-davinvi-003 and chatGPT, both initially refused to spell the garbage word "khfsihciweb" but after insisting they did it. It can even invent readings for acronyms that don't exist: Acronym: CNN Expand: Convolutional Neural Network Acronym: PNP Expand: Plug and Play Acronym: SHDJ Expand: Super High-Definition Jukebox Acronym: HEDUE Expand: Higher Education Digital User Experience Acronym: SHADBEESHQEKSIHXHGESLS Expand: Super Highly Advanced Digitalized Bionic Electronic Exoskeleton Super Human Quality Exoskeleton Kit for Super High-Gravity Environments and Space Living
- nebulous1 4y agoYou can ask it to spell out words, which it will do. I asked it to do this and it was successful. I then intentionally misspelled a word (e-i transposition), and it gave me the correct spelling. I asked it if that was how I spelled it, and it confirmed that it wasn't but that the way I had spelled it was incorrect (and noted the e-i transposition). So while it may may receive input as word tokens, I'm not sure that that means it can't "see that at all". I will note that I don't have a good understanding of chatgpt at all.
- decremental 4y agohttps://beta.openai.com/tokenizer https://beta.openai.com/tokenizer Test with: 6f6b25515e18b4c41927e2dce3ffc589 You'll see that it spits out single character tokens.
- dizhn 4y agoBut it later makes the claim that 7, 8, 9 sounds like "scared". ?