6 ms·
That they can't do this sort of simple question speaks volumes to the entire approach. I don't think generative AI will ever be able to reach AGI, and most peo
by NoCoooode 2y ago
That they can't do this sort of simple question speaks volumes to the entire approach.
I don't think generative AI will ever be able to reach AGI, and most people selling LLM today pretend it is AGI
- GaggiX 2y agoThe fact that LLMs are usually trained on tokens and not on characters, doesn't really speak about what generative AI is going to reach or not. >most people selling LLM today pretend it is AGI Who are these "most people"?
- smokedetector1 2y agoELI5 why are tokens not a single letter?
- WhitneyLand 2y agoSuch an architecture could be implemented, it could use one token per letter, or one token per word, instead of the typical 0.75 per word we see. The choice just comes with trade-offs in memory usage, compute, and effectiveness of the model in various scenarios. So what we ended up with was a pragmatic/engineering decision rather than a theoretical or fundamental constraint.
- alach11 2y agoAll it speaks to is that tokenization is weird and introduces artifacts to LLM performance. Counting letters is a trivial task when you're staring at words on a screen. It's much harder when you're perceiving vectors based on parts of words. The fact that LLMs find certain things easier/harder than humans is completely unsurprising, and there are much more interesting benchmarks to use to compare one LLM to another.
- jrflowers 2y agoThis is a good point. While LLMs being incapable of reliably doing a simple task that’s been doable by computers since the punch card days is an important consideration for anyone that might be thinking about using them for anything other than as a toy, this fact is uninteresting because of Reasons
- space_fountain 2y agoLLMs can clearly solve problems that computers up to now couldn't. They can't solve all problems and this should definitely be a cautionary note to anyone who wants to use them as an artificial general intelligence, but this take seems no different to someone looking at a punchcard computer and going, it can't even recognize typos or categorize images, what good is this? We've already had human computers who can do everything these can do, and can recognize images and notice typos
- evilduck 2y agoAlso humans would revert to explicitly using an algorithm and external storage like a sheet of paper with tally marks or a spreadsheet or even a computer program if you scale the question up to a full sheet of text or a whole book or a collection of books (we probably do it at a single word size too, but it's more intuitive than explicit behavior for most folks when the count sum is around 8 or less). LLMs can't effectively execute algorithms similarly in their context, nor can they memorize new data or facts it was given without providing it tools like function calling or embeddings. If you give LLMs tool calling and storage mechanisms then counting letters in words becomes pretty damn reliable.
- jrflowers 2y ago> going, it can't even recognize typos or categorize images, what good is this? No one said that LLMs aren’t good for anything. I pointed out — in response to another poster downplaying mention of a well-known and undisputed limitation that LLMs often have — that it is valid to consider these well-known and undisputed limitations if one is considering using them for anything other than a toy. It is downright silly to discourage discussion of well-known and undisputed limitations! The only reason for that can only be entirely emotional as there is genuinely nothing tangible to be gained by being steadfast in silence about a fact that isn’t up for debate.
- doctorpangloss 2y agoCounting shit, like cells, peaks in signals, people, inventory, fingers, and votes, is hard, tedious and important to business and life, so I don’t know dude, it seems like a great benchmark to me. Countless posts wasted on denying this simple and obvious fact.
- deleted 2y ago[deleted]
- BoorishBears 2y agoIt's like using a hammer to turn a screw and calling it useless. To envision what a next generation model bound by the same constraints should do, it'd be to recognize that it can't count tokens and use code access to write code that solves the strawberry problem without prompting. Asked to count cells it'd be a model that could write and execute OpenCV tasks. Or to go a step further, be a multimodal model that can synthesize 10000 varations of the target cell, and finetune a model like YOLO on it autonomously. I find arguments that reduce LLMs to "It can't do the simple thing!!!!" come from people unable to apply lateral thinking to how a task can be solved.
- deleted 2y ago[deleted]
- doctorpangloss 2y ago> To envision what a next generation model bound by the same constraints should do, it'd be to recognize that it can't count tokens and use code access to write code that solves the strawberry problem without prompting. The VQA problems I'm describing can be solved seemingly in one case but not combined with counting. Counting is fundamentally challenging for sort of unknown reasons, or perhaps known to the very best labs who are trying to tackle it directly. Another POV is that the stuff you are describing is in some sense so obvious that it has been tried, no?
- BoorishBears 2y agoI don't get what you mean by "unknown reasons", we understand that counting tokens requires a type of introspection transformer models can't do while operating on tokens. What I described is tried, and works, but the models are still not cheap/fast/reliable enough to always do what I described for every query. The difference between what I described and directly asking the model to count is that we know the models can get cheaper, faster, and more reliable at what I described without any earth shattering discoveries Like I don't see any reason why GPT 10 will ever be able to count how many letters there are in the word strawberry without a complete paradigm shift in model building... but going from GPT 3 to GPT 4 we already got a model that can always write the dead simple code required to count it out, and the models that can do so are already getting cheaper and faster every few months without any crazy discoveries.
- throw101010 2y ago> most people selling LLM today pretend it is AGI Who exactly does this in this space? Would be good to be able to call them out on it right now.
- swyx 2y agoimagine being so confidently wrong about AI
- jimbokun 2y agoIn isolation, probably not. But it's likely to be an important component in an AGI system. I suppose the interesting question is how to integrate LLMs with more traditional logic and planning systems.
- bondarchuk 2y agoFor all I care we will have superhuman AGI that still can't count the Rs in strawberry. Some humans are dyslexic and all are subject to weird perceptual illusions; doesn't make them any less human-level intelligent.
- InsideOutSanta 2y agoIn my opinion, the problem with the strawberry question is that it is both a bad example because you don't need an LLM to count the number of r's in a word, and it's a bad measure of an LLM's capabilities because it's a type of question that all LLMs are currently bad at. Having said that, the 40b model wasn't able to answer any of my real-world example questions correctly. Some of these (e.g. "how do I add a sequential number after my titles in an HTML page using just CSS, without changing the page") are questions that even some of the better small local models can answer correctly. It gave very authoritatively sounding wrong answers.
- theGnuMe 2y agoGodel's Strawberry.