5 ms·
The issue with this paradigm (see also: Hutter Prize) was its insistence on lossless compression. Intelligence is just as much about knowing what to throw away
by optimalsolver 2y ago
The issue with this paradigm (see also: Hutter Prize) was its insistence on lossless compression.
Intelligence is just as much about knowing what to throw away as what to keep. Any nontrivial cognitive system operating in a nontrivial physical environment will necessarily have a lossy model of the world it's embedded in.
Also of interest, compressionism, a theory of mind based on data compression:
https://ceur-ws.org/Vol-1419/paper0045.pdf https://ceur-ws.org/Vol-1419/paper0045.pdf
- theendisney 2y agoYou could throw away 95% of enwiki without losing anything of value. This sounds like a joke but it would make the result worth reading which sounds like it is worth the exercise.
- sfink 2y agoI tend to agree, but is that really a fatal flaw? A lossy compression scheme that gets it 90% right only has to encode the delta for the remaining 10%. Which is a big cost, sure, but the alternative is evaluating how important the thrown-away stuff is, and that evaluation is rife with subjective value judgements. There are no right answers, only differing flavors of wrong ones. It's the question of what is important to generalize over, and nobody is ever going to agree on that.
- sdenton4 2y agoYou can turn any lossy compression scheme into a lossless scheme by encoding the error... The lossy scheme should be aiming for low error, making the error encoding progressively smaller as the lossy scheme improves. What's harder to deal with from a measurement perspective is sematic equivalence. calling some kinds of errors zero-cost, but not having a great way to categorize what the loss of, exactly. But it's kinda what you want for really extreme compression: the content is equivalent at a high level, but may be a very different byte stream.
- igorkraw 2y agoGroup theory is the abstraction you want
- jebarker 2y agoCan you explain? I'm struggling to see how group theory is relevant here
- igorkraw 2y agoThis part >What's harder to deal with from a measurement perspective is sematic equivalence. calling some kinds of errors zero-cost, but not having a great way to categorize what the loss of, exactly. But it's kinda what you want for really extreme compression: the content is equivalent at a high level, but may be a very different byte stream. Is basically saying > What's harder is defining a reconstruction process in terms of a "semantic group" i.e. an output encoding and associated group actions under which the loss is invariant, and having the group actions express the concept of two non-identical outputs being"equivalent for the purposes of downstream processing". Taco Cohen is one of the pioneers of this line of research, and invariant, equivariant and approximately iv/ev architectures are a big thing in scientific and small data ml
- sdenton4 2y ago"an output encoding and associated group actions under which the loss is invariant" What loss is that, exactly? One of the difficulties in speech processing is that we generally don't have a great model for speech quality or equivalence. The human hearing system is a finicky and particular beast, and furthermore varies from beast to beast, depending both on physiology and culture. Good measures (eg, Visqol, a learned speech quality metric) tend to be helpful for measuring progress when iterating on a single system, but can give strange results when comparing different systems. So it 's easy to imagine (say) pushing the generated speech into some representation space and measuring nearness in that space (either absolute or modulo a group action), but it begs the question of whether nearness in that space really represents semantic equivalence, and how to go about constructing it in the first place. Let alone why one would bother allowing some group symmetries into the representation when we plan to define a loss invariant under those symmetries... Throwing group theory at speech representations feels like a solution in search of a problem, as someone who has worked a lot with group theory and speech compression.
- vintermann 2y ago> Intelligence is just as much about knowing what to throw away as what to keep. Sure, but you can see it as two steps. Decide what doesn't matter at all, and throw it away. Then compress the rest losslessly (according to how surprising it is, basically the only way) I think a theory of mind based on data compression is backwards, for this reason. When you have "data", you've already decided what's important. Every time you combine "A" and "B" into a set, you have already decided 1. That they are similar in some way you care about 2. That they're distinct in some way you care about (otherwise, it would be the same element!) ... and you do this every time you even add two numbers, or do anything remotely more interesting with "data". There is no "data" without making a-scientific statements about what's important, what's interesting, what's meaningful, what matters etc.
- abecedarius 2y agoYour issue applies equally to GPT-3 (and afaik its successors): the log likelihood loss which the GPT base model was trained on is exactly the "lossless compression" length of the training set (up to trivial arithmetic-coding overhead). The question I'd raise is why the route that got traction used more ad-hoc regularization schemes instead of the (decompressor + compressed) length from this contest and the book I linked.