7 ms·
Compressing Hegel (2020)
- TurkishPoptart 3y agoThis piece could benefit from some H1/H2 headers to break up the text.
- readthenotes1 3y agoOr paragraph breaks. Tldr: tldr.
- eyelidlessness 3y agoIt has Hegel right there in the header, it’s self-compressing to whatever your interest threshold happens to be.
- AnotherGoodName 3y agoMore directly they are both about prediction. If I can accurately predict the next X bits I don't need to store those. If I can accurately predict the next X bits of information that will happen after a decision I can make a reasoned decision. That's it. That's the link as plainly as possible. If you want more detail read up on arithmetic coding and how it can take predictions as in input to data compression and then Marcus Hutters papers on general purpose AI requiring the exact same thing.
- motohagiography 3y agoYears ago at the office I had a conversation about complexity where I asked someome who knew more than me whether feature extraction, similarity metrics, kolmolgorov complexity, and graph isomorphism were really all just the same information encoding and compression problems, and he said something like, "well, sure, but that means everything is just a compression problem" which I thought was pretty funny. This author seems caught in a similar loop. On a related epistemological loop - I was just using the openai tool to query some of Kripke's modal logic ideas, and find whether they could be expressed in code (they can, for a given theorem in modal logic, gpt can represent the theorems in lisp if you want), and what struck me about the whole premise was that he invented not so much a broader logic that was revelatory or quantitatively truth finding - but a neutralizing, critical logic designed to dilute others. Gödel's theorems are apparently for logics that can produce arithmetic, which seems basic an necessary, but Kripke figured, "fine, what about logics that don't produce arithmetic?" (the result is basically digraphs.) Apparently there's nothing Kripke produced that can't be produced symbolically with digraphs, and like the author's compression example, it's just digraphs all the way down. Same flaw. Kripke's ideas seem like fun philosophical ideas, but they're more of a scheme backfitted to a familiar ideology, and then you see his ideas come up often again as foundations for other less rigorous and less quantitative theories. If you've ever heard people use the word "modality," they're attempting to control a discussion using some of these same techniqes in what they percieve as a power struggle without any need for alignment to truth or reality. By adding parallel logical systems without the same criteria for consistency, Kripke created a tool for infinite uncertainty whose main feature is to neutralize the concept of logical truth in its subjects. I think it will be immensely useful in creating AI's, but it will also be the basis for the truly dangerous parts of it. Modal logic does not relate to the universe or need external consistency, it only exists as a kind of solvent for another existing logic it effectively criticizes. If there were such a thing, I think it may be a recipe for actual evil. Anyway, good luck with the "everything is compression" thing, and even the basilisk that is Kripke's digraphs dressed up as ontology. I don't think there is a there there, but if we get to talk about this stuff, there may still be some fun to be had.
- mrwnmonm 3y agoLove this comment xD https://wittgenfine.substack.com/p/compressing-hegel/comment/3979023 https://wittgenfine.substack.com/p/compressing-hegel/comment...
- kelseyfrog 3y ago> Lossless compression is equivalent to intelligence. The article has all of the foundations to make a much more reasoned and perhaps more interesting claim: intelligence is equivalent to the size of one's abstraction inventory, and abstraction utility is measured by prediction success in abstraction-space. The fact is that humans' short term memory capacity[1], necessitates the hoisting of raw sense data into an abstract space, and doing so implies a very real lossy process. We don't really make much of a fuss when ontologizing the world, but it's a bit strange doing that specifically when referencing Hegel. The author is as guilty as anyone who, during the process of abstraction, forgets they have done so and thinks that the concept of a cat implies that cats are some universal essence made manifest. Forms are constantly trying to convince us that they're real. 1. https://en.wikipedia.org/wiki/The_Magical_Number_Seven,_Plus_or_Minus_Two https://en.wikipedia.org/wiki/The_Magical_Number_Seven,_Plus...
- rektide 3y agoHoisting and... And applying/recalling. Quite a lovely post, thank you, well constructed. One other theme that shows up (briefly) is that storing the data isn't good enough. Being able to access your compressed data, being able to apply yiur compressed models to new applications is, in my view, another core aspect of intelligence, one that requires & uses real ingenuity. Being a highly associative person is, IMO, a colossal form of intelligence & intellect.
- voidhorse 3y ago> The author is as guilty as anyone who, during the process of abstraction, forgets they have done so and thinks that the concept of a cat implies that cats are some universal essence made manifest. Forms are constantly trying to convince us that they're real. The later Marxists had a term for this, reification, and I'm convinced it is probably the most common fallacy human beings make.
- kelseyfrog 3y agoUsing the technical term has tended to enrage folks here before. "Why should I have to read The Social Construction of Reality to understand what you've written," has the same sensibility as "Why should I have to know a programing language to write programs?!" But yes, that's precisely it.
- kaibee 3y agoIts worth noting that random noise is also incompressible.
- voidhorse 3y agoIt's ironic that an author who takes their moniker after Wittgenstein would espouse a view that Wittgenstein's own later work (the investigations) practically argues directly against. The perspective presented here is soaked through and through with information theoretic bias and clearly stems from an overly digital consideration of human experience. "Intelligence" is intimately related to context, use, community, and goals. We don't call someone intelligent because we can point to some lossless compression algorithm implemented by his neurons, we call him intelligent when he produces the behaviors we desire in a given situation. Consider for instance, how this theory fails to account for athletic knowledge. It seems fair to state that people with tangible skills requiring the use of the body posses some kind of intelligence, it seems less accurate to try and describe this form of intelligence as lossless compression. The compression metaphor really only works for intelligence with respect to the manipulation of symbolic representations and even then I think there are plenty of counter examples of what we'd call intelligent behavior that would suggest not all cases are reducible to "lossless compression". It is all highly dependent on the questions you ask, as Wittgenstein well knew. Personally, I do not feel this is a good piece of philosophy and it's symptomatic of a trend toward "computormorphization" which has been rampant since the computer emerged (it's like anthropomorphism except interpreting the behavior of other (actually living things!) as though they behaved just like computers). You can also see this in the reification of concepts like information (we talk about information as though it were some objective material object, but information is not a substance--only interpreters produce information; a text is a vehicle for information but it does not "contain information"--the information is produced by the reader of the text and fully depends on what distinctions they deem relevant)
- canjobear 3y agoI don't think the information-theoretic compression-based view is at all incompatible with Wittgenstein in Philosophical Investigations, although it's absolutely against the view in the Tractatus. In the Tractatus, something is meaningful when it is isomorphic to some structured symbolic representation. Information theory is based on a rejection of that kind of view (from Shannon 1948): "The fundamental problem of communication is that of reproducing at one point either exactly or approximately a message selected at another point. Frequently the messages have meaning; that is they refer to or are correlated according to some system with certain physical or conceptual entities. These semantic aspects of communication are irrelevant to the engineering problem. The significant aspect is that the actual message is one selected from a set of possible messages." In other words, information has nothing to do with isomorphism to a structured representation with definite relationships with the world, nothing to do with symbolic representations or any other kind of "representations". It's simply the ability to predict one thing from another thing. Shannon calls these things "messages" but mathematically the logic applies to any set: for example we can analyze an agent's policy information-theoretically, in terms of how much information about the agent's input state is contained in its output actions. I think this kind of view is highly compatible with later Wittgenstein, for whom "meaning" (if it's anything) is an abstraction of how language is used to coordinate joint action---that is, situations where people are trying to predict and control each others' behavior. To address your examples, > We don't call someone intelligent because we can point to some lossless compression algorithm implemented by his neurons, we call him intelligent when he produces the behaviors we desire in a given situation. Suppose we give someone input A and he outputs desired behavior f(A), and then we give him input B and he outputs desired behavior f(B). If we give him input C will he output the desired f(C)? Only if he has learned to predict what f(C) should be as a function of C and the training data, that is, if he has learned to predict what we consider the "desired" behavior to be. In that sense, intelligence here requires prediction, which is exactly the same thing as compression. > It seems fair to state that people with tangible skills requiring the use of the body posses some kind of intelligence, it seems less accurate to try and describe this form of intelligence as lossless compression. Again, the agent's problem here is to select the right motor action given sensory input. Suppose you've seen a tennis ball coming in with velocity vector x and you know how to hit it to win the game, and you've seen a tennis ball coming in with vector y and you know how to hit it to win the game. What will you do with vector z? If your policy is good, it will put high probability on only actions that hit the ball in a way that makes you win. And if your policy is putting high probability on that action, that's mathematically the same as saying it provides an efficient compression of that action (high probability = small number of bits in an encoding). This is very abstract but once you adopt this view it gives you a lot of very useful conceptual tools. For example, you can talk about the "channel capacity" of a policy in terms of how many bits of information it can "transmit" from input states to output actions, and this channel capacity turns out to be a very intuitive measure of the complexity of the policy that you can use to analyze human behavior (one example, [1]). [1] https://gershmanlab.com/pubs/GershmanLai21.pdf https://gershmanlab.com/pubs/GershmanLai21.pdf
- evertedsphere 3y agoreminds me of https://arxiv.org/abs/0812.4360 https://arxiv.org/abs/0812.4360 > I argue that data becomes temporarily interesting by itself to some self-improving, but computationally limited, subjective observer once he learns to predict or compress the data in a better way, thus making it subjectively simpler and more beautiful. Curiosity is the desire to create or discover more non-random, non-arbitrary, regular data that is novel and surprising not in the traditional sense of Boltzmann and Shannon but in the sense that it allows for compression progress because its regularity was not yet known. This drive maximizes interestingness, the first derivative of subjective beauty or compressibility, that is, the steepness of the learning curve. It motivates exploring infants, pure mathematicians, composers, artists, dancers, comedians, yourself, and (since 1990) artificial systems.