8 ms·
>Figure 1: We prove that SSMs, like transformers, cannot solve inherently sequential problems like permutation com- position (S5), which lies at the heart of st
by optimalsolver 2y ago
>Figure 1: We prove that SSMs, like transformers, cannot
solve inherently sequential problems like permutation com-
position (S5), which lies at the heart of state-tracking prob-
lems like tracking chess moves in source-target notation
(see Section 3.2), evaluating Python code, or entity tracking.
Thus, SSMs cannot, in general, solve these problems either
Do Microsoft & friends who are about to build trillion dollar AI data centers know about these proven limitations of transformer-based architectures?
- mjburgess 2y agoYes. The solution is to hoover up all the queries people are currently putting to ChatGPT, save how people have "prompt engineer"ed the solution (ie., answered the question themselves); and then hope that you can just feed this back to people without them noticing. The only open question is whether people have common-enough queries for this charade to work out. It seems there's quite a lot at least. But this number will decrease over time for various reasons. So it's a game of building a system that can be retrained on the answers people are giving it fast enough that people don't notice where the answers are coming from.
- FeepingCreature 2y agoCounterpoint: the more networks internalize the patterns behind engineered prompts, the closer they get to being general reasoners.
- mjburgess 2y agoStatistical patterns in text tokens have nothing to do with reasoning. They lack such a property to learn in the first place. This would be more obvious, I guess, if the tokens were in an alien language. Consider a translation to 1,000 different alien languages of any given novel.. there is no distributional property they would share. Statistical AI is just a way of sampling from a historical dataset with a similarity metric. It only works to answer questions if you're sampling from a (Q, A) database in the same language the user already understands. The question was answered by a reasoner, it is now answered by a system which replays answers.
- FeepingCreature 2y agoThis is just chinese room. But also, I disagree that the languages would not share properties. A novel is too small. Consider a network trained on GPT-3 scale datasets in 1000 alien languages. The shared structures behind the sentences will be the same, even if the grammar is completely different. Stars will be stars, moons will be moons. I'd bet you that the model would be able to translate shared concepts between those languages even if it had not been trained on the same novels, same as GPT can translate terms that it has not seen in dictionary pairs. (I have no idea if it can! But I'm confident enough that I'm willing to just say it can, at risk of being proven wrong. If GPT can't do that, I'm fundamentally misunderstanding how it works - that is, I don't have a paper offhand showing that GPT shares concept neurons between languages, but I'm willing to bet I could find one if I went looking.) In other words, if you co-trained GPT-4 on Earth Internet and Alien Internet, there's a good chance it'd end up able to translate English to Alienese, purely as an emergent ability, if it had the concept of translation at all. Intelligence is compression. With sufficient abstraction (layers) and sufficient volume (dataset), any description of the same reality will assume the same structure. And no learning algo worth its salt will keep two identical structures around.
- mjburgess 2y agoAlmost every predicate in natural languages is an arbitrary association of properties in the world. There isn't any property a person has "bald", nor are there "tree"s in the world. What bundles of properties the vast majority of language names are a product of historical and contingent associations we've made for practical reasons. Likewise, languages do not have the same distributional structure. There is no reason an alien language would use discrete tokens to name properties, nor co-locate tokens by linear position in a 'sentence'. Historical human languages did not; using, eg., the full 2D structure of the clay tablet. To suppose that it is the glyphs and their colocation which somehow bare meaning is a nonesensical superstition. The world is what our words mean, and it is we, in that world, who provide them their meaning. We can do so with arbitrary linguistic structures.
- FeepingCreature 2y ago
- empath75 2y agoNot being able to solve a problem "in general" does not mean that it can't solve specific and useful instances of the problem. There are lots of SAT solvers that can't solve SAT problems in general (in any reasonable amount of time), but nevertheless can solve many useful categories of SAT problems.
- canjobear 2y agoPeople at Microsoft Research certainly do. But the bet is that, with scale, these limitations won't matter in practice. For example Transformers can't recognize well-nested brackets (like {[()]}) to infinite depth; the depth that Transformers can recognize is limited by how many self-attention layers they have. But in practice, you rarely need much depth.
- toxik 2y agoNotably, humans also cannot track this to infinite depth. Somehow we know this and reach for algorithmic solutions (like pen and paper).
- logicchains 2y agoInterestingly LLMs also do better with pen and paper reasoning: https://arxiv.org/abs/2310.07923 https://arxiv.org/abs/2310.07923 .