7 ms·
Congrats! I went to your demo and asked for words that end in agi. This is what I got: -- agi, agi, agi, agi, agi, agi, agi These are some of the words that
by primordialsoup 3y ago
Congrats! I went to your demo and asked for words that end in agi. This is what I got:
--
agi, agi, agi, agi, agi, agi, agi
These are some of the words that end in agi. You can also use the word agi in a sentence. For example, "I am going to the grocery store to get some agi."
These are some of words that end in agi.
These are some words that end in agi.
maximize, maximize, maximize, maximize, maximize, maximize, maximize, maximize
These are some words that ends in agi
--
So I think this needs more work to get to "as good as ChatGPT". But having said that, congrats on the landing
- avereveard 3y agoyeah as usual these model can barely sustain a conversation and fall apart the moment actual instructions are given. typical prompt they fail to udnerstand: "what is pistacchio? explain the question, not the answer." all these toy llm: "pistacchio is..." gpt is the only one that consistently understand these instructions: "The question "what is pistachio?" is asking for an explanation or description of the food item..." this makes these llm basically useless for obtaining anything but hallucinated data.
- vidarh 3y agoIt only makes them useless.of you insist on asking them in ways you already know will provide bad results instead of adapting your prompts. This is a bit like complaining that your compiler refuses to produce the right outputs for code you've already determined is incorrect.
- avereveard 3y agoAsking LLM from things they learned in training mostly result in hallucinations and in general makes you unable to detect by which amount they are hallucinating: these models are unable to reflect on their output, and average output token probability is a lousy proxy for confidence scoring their results. On the other hand, no amount of prompt engineering seems to make these LLM able to do question and answer over source documents which is the only realistic way by which factual information can be retrieved You're welcome to bring examples of it tho if you're so confident.
- vidarh 3y agoI've had ChatGPT build a fire nctioning website, write a DNA server, fill in significant portions of specs, all without the problems you describe. I'm never going back to doing things from scratch - it's saving me immense amounts of time every single day. The only reasonable conclusion is that the way you're promoting it is counterproductive.
- avereveard 3y agoGood thing then that I specifically mentioned gpt as being able to follow instruction and that I was specifically mentioning the other models. You're welcome to demonstrate the same ability on other models tho.
- vidarh 3y agoYou can get useful results out of a whole lot of them as long as you actually prompt them in a way suitable for the models. The point I made originally was that if you just feed them an ambiguous question, then sure, you will get extremely variable and mostly useless results out. Ironically, And I mentioned ChatGPT because from context of your comments here it was unclear on first read-through what you meant. Maybe consider that it's possible your prompting is not geared for the models you've tried. Not least, specifically given that if you expect a model to know how to follow instructions, when most of them have not been through RLHF you're using them wrong. A lot of them needs prompt shaped as a completion, not a conversation.
- avereveard 3y agoyou're welcome to provide examples to prove your points.
- vidarh 3y agoI have nothing to gain from spending time testing models for you because whatever I pick will just seem like cherry picking to you, and it doesn't matter to me whether or not you agree on the usability of these models. They work for me, and that's all that matters to me. Try a a few completions instead of a question. Or don't
- brookst 3y agoIt's a fair criticism, and ChatGPT does better, but this isn't a great test of model quality. All LLMS that rely on tokenization struggle with being introspective on language. Try asking chatGPT to count how many e's are in a sentence, or to list all words that start with "to" and end wide "de". I haven't heard anyone describe the phenomenon clearly, but I expect it is a challenge with reasoning over both intent of the prompt and specific token IDs.
- mckirk 3y agoThe problem is that these models do not have any working memory they could use to carry out such tasks, which are on a meta-level when seen from a language perspective. They can only go with their 'gut instinct' for selecting the next word, they can't 'consider and ponder the problem internally' first.
- jameslevy 3y agoPerhaps prompting the model to first describe its approach to answering the question. This type of chain-of-thought technique can yield better results.
- sp332 3y agoThe problem is that the input is tokenized before the model gets it as input. It does not see the individual letters "t" + "o". It gets one single token, #1462. The word "toe" is another single token, #44579. Maybe over time it could learn from context that inputs that start with #44579 also satisfy the constraint of starting with #1462, but that's a lot of work and it's not going to happen for all combinations of letters.
- chaxor 3y agoIt's not a character based model (likely - although it's closed source so anything is technically possible behind the scenes) so this makes some sense. The system can infer some relationships, which may be why 'agy' is conflated with 'agi' interestingly, but the tokenization process yields sequences of 'symbols' or indexes that are decided to English - so the system has a more difficult task when asked about 'e's (probably something like token 4893) and has to determine which tokens (e.g. [358,284840, 58292, 4830104, 57282, 4829193, 58282, 384, 24945] contain 'e's or token 4893). None of them do directly it seems - but 58292 may be 'ee' - so you would get this wrong as well.
- TwoFactor 3y agoThats an interesting test. Here's what I got from ChatGPT: ---GPT-3.5--- Here are some words that end in "agi": Strategy Swarajya Arthroplasty Sialagogue Podagric Gynecology Physiognomy Ophthalmology Esophagitis Otalgia --- GPT-4 --- Here are some words that end in "agi": Swaggy Raggi Magi Gagi Stagi Please note that some of these words may not be commonly used or may be specific to certain dialects or regions.
- armchairhacker 3y agoStagi isn't a word (unless you count Lojban). Gagi isn't a word unless you could Filipino slang.
- mejutoco 3y agoTo be fair the question did not specify the language and included a disclaimer about it.
- ksaj 3y agoEven if you don't consider an Italian word as a word: It's a last name. It's a brand name. It is several companies' name. It belongs in the list just fine.
- mejutoco 3y agoIt seems you agree with me, I do not understand. Wrong thread maybe you replied to?
- ksaj 3y agoYes. That was agreeing.
- littlestymaar 3y agoThen the word rhamanagagi (which I just made up) is a word that would technically belong to the list just fine, it definitely not answered to the implicit intent of the question. The strength of LLM is their ability to answer to unprecisely specified questions, being able to guess the speaker's intent, but in this particular case, it's failing the test.