5 ms·
An interesting example of this is: There are 6 “a”s in the sentence: “How many ‘a’ in this sentence?” https://chatgpt.com/share/677582a9-45fc-8003-8114-edd2e6
by frikskit 2y ago
An interesting example of this is:
There are 6 “a”s in the sentence:
“How many ‘a’ in this sentence?”
https://chatgpt.com/share/677582a9-45fc-8003-8114-edd2e6efa289 https://chatgpt.com/share/677582a9-45fc-8003-8114-edd2e6efa2...
Whereas the typical “strawberry” variant is now correct.
There are 3 “r”s in the word “strawberry.”
Clearly the lesson wasn’t learned, the model was just trained on people highlighting this failure case.
- bwfan123 2y agoReminds me of software i have built which had some basic foundational problems. Each bug was fixed with a data-patch that fixed the symptom but not the cause. hence we continually played whack-a-mole with bugs. we would squash one bug, and another one would appear. same with llms, squash one problem with a data-fix, and another one pops-up.
- sealeck 2y agoIt also fails on things that aren't actual words For example, the output for "how many x's are there in xaaax" is 3. https://chatgpt.com/share/677591fe-aa58-800e-9e7a-81870387bebf https://chatgpt.com/share/677591fe-aa58-800e-9e7a-81870387be...
- deleted 2y ago[deleted]
- Isinlor 2y agoTransformers are very bad at counting in one feed forward pass, you need to explicitly tell them to use a counter in autoregressive fashion like here: https://chatgpt.com/share/6775cb37-4198-8007-82cb-e897220827f8 https://chatgpt.com/share/6775cb37-4198-8007-82cb-e897220827...
- deleted 2y ago[deleted]
- Isinlor 2y agoTransformers are very bad at counting due to how their internals work. But if you ask them to use explicit counter the problem disappears: https://chatgpt.com/share/6775c9a6-8cec-8007-b709-3431e7a2b24e https://chatgpt.com/share/6775c9a6-8cec-8007-b709-3431e7a2b2... Basically one feed forward is not Turing complete, but autoregressive (feeding previous output back into itself) are Turing complete.
- frikskit 2y agoThis makes it worse IMO. I was starting to think it didn’t have a letter by letter representation of the tokens. It does. In which case the fact it didn’t decide to use it speaks even more towards its unsophistication. Regardless, I’d love if you would explain a bit more why the transformer internals make this problem so difficult?
- Isinlor 2y agoWhen Can Transformers Count to n? https://arxiv.org/html/2407.15160v2 https://arxiv.org/html/2407.15160v2 The Expressive Power of Transformers with Chain of Thought https://arxiv.org/html/2310.07923v5 https://arxiv.org/html/2310.07923v5 Transformer needs to retrieve letters per each token while forced to keep internal representation still aligned in length with the base tokens (each token also has finite embedding, while made out of multiple letters), and then it needs to count the letters within misaligned representation. Autoregressive mode completely alleviate the problem as it can align its internal representation with the letters and it can just keep explicit sequential count. BTW - humans also can't count without resorting to sequential process.
- frikskit 2y agoThanks!