5 ms·
Notes on a New Philosophy of Empirical Science (2011)
- mjburgess 3y agoScientific models are causal. Compression is a condition on association. I don't really know what more needs to be said here.
- tgv 3y agoIt's an idea about how to judge models. A model's predictive capabilities are modelled (indeed) as compression, like "how many bits do you need to set up your model and correct its output". It might be nice to compare fairly complete models on a well defined domain, but I can't see it as a general guiding principle. It would get theorizing stuck in a local minimum.
- mjburgess 3y agoScience isn't interested in this kind of prediction. That's just engineering. Causal models give counter-factual predictions for existence claims (eg., that a planet exists because the orbit of two other planets doesn't follow the causal model). Science, in most cases, prefers models with poor "engineering predictions" (ie., point estimates of observables) because they have vastly superior explanatory power. In most cases it would be a catastrophe for a scientific model to be making good estimates of observables, because we know a priori, that observables aren't fully determined by the model (eg., just consider that F=GMm/r^2 basically didnt apply to most observations of the solar system when it was formulated by newton; nor really does it much today). Explanatory power is not a property of compression, nor association, nor "prediction" in this engineering sense. Consider here that a lossless model of the solar system would never have yielded newton's law of graviton (since most of the objects in the solar system are unknown). This entire project is just, "what if science were like ML?" -- an interesting question only because how vast the gap is; and how absurd the suggestion.
- dr_dshiv 3y agoHow can you tell the difference between a causal model and a predictive model? Isn’t it just the elegance of the model? And isn’t elegance just succinctness? Because another approach could focus on human comprehensibility. Indeed, a leading theory of scientific progress focuses on the advancement of noetic understanding. But, then we should really be irritated by things like quantum mechanics and love things like nutrition (we might get poor predictions but good understanding). I’m not sure myself.
- mjburgess 3y agoNo, it's a completely different semantics. F = GMm/r^2 does not mean anything about {P(F|M), P(F|m), ..) -- this this the semantics of associative statistical models, as used in ML. Rather it means a gravitational force is caused by the interaction of two masses over a distance r. Here `F` refers to a force via a scientific model, etc. The formula is a short-hand conceuqnece of a family of explanatory models about mass, inertia, gravity, forces, etc. and is only valid when used in the context of those models. Eg., you cannot equate F = GMm/r^2 to F = kQq/r^2. Not least since we have no gravitational model which applies to tiny charged particles. The formula used in scientific modelling do not have either mathematical or statistical semantics (ie., they dont refer to numbers, nor to associations). They refer to bits of the world via explanations; and are only valid insofar as these explanations apply. Though science seems obsessed by formula, in terms of the goal of science, it's the least important part of the scientific model. Explanations are the goal, and highly circumstantial consequences of these are given formula partly for illustration, partly for engineering. There's rarely anything in any dataset whose model would even be useful for building an explanation. Explanations are build via counter-factuals, and these are resolved by experimental data -- they are not made by it. Indeed, in many cases, the experimental data would be no where nearly modelled lossessly by the candidate hypothesis.
- dr_dshiv 3y agoHere’s a toy example based on a project I’d like to do. I love faraday waves and Chladni plates — roughly speaking, cymatics. Now, if I record high res video of the waves on water in a dish, vibrating at different frequencies, I could likely create a diffusion model that had some latent “understanding” of the relationship between the proportional frequencies of sound and the waves in a particular sized dish. I could test this by holding out certain frequencies and observing whether the diffusion model could recreate them. So, I can ask the question of whether the AI model learned wave physics. Here, there would be no formula, per se, and no explanation. Merely a computational model that could make predictions about physical phenomena. Now, what makes this less scientific than a formula and set of explanations for the experiments? Is it because I can relate the the explanations linguistically to everything else I know about science? So, in that case, If I jointly developed a language model to do the same thing as the diffusion model, ie make predictions based on data—but linguistically capable of connecting the outcomes to scientific concepts, then would then it be scientific?
- Loquebantur 3y agoScientific models clearly represent a compression of measurement data? Scientific models aren't necessarily "causal" to begin with. They are functions that give predictions about measurements. It is these predictions that are tested against data, not the function's confabulated "causal" justifications. People learn from data not adhering to predictions. This difference from model functions can be compressed, if not random. This compression then might reflect in the form of modularization in justifications, which again is interpreted as causal relationships.
- mjburgess 3y ago> Scientific models clearly represent a compression of measurement data? Nope! No theory of heat is a compression of therometer readings; no theory of gravity, of orbits; no theory of atoms of spectra. No theories compress measurments! Such a thing is pure superstition. Heat is not the motion of thermometer fluid. > not the function's confabulated "causal" justifications. Nope! We construct an experiment by counter-factual analysis of its causal semantics; we do not simply test whether observable quantities match prior data. Arbitary associative models match arbitary amounts of prior data. This is the opposite of science. We test scientific models by creating new experiments; it isnt "the data" which matters here, but that the experiment is designed to test the causal assumptions of the model. If the experiment doesn't: control causes, identify novel measures with potential causes, etc. then any data collected is useless. This is why you need, you know: randomised controlled trials, microscopes, satellites, ... etc. "Data" in the ML sense does not matter. This is pure superstitious pseudoscience. Science is a process of creating data under experimental conditions designed to be counter-factual tests of theories. Science is about the data generating process (reality), not our measurements of it.
- Loquebantur 3y agoI'm sorry, but I think you misinterpret what compression is all about? Heat is a random process. That process has a non-random component though. The phenomenon is compressed by describing it as a random distribution of impulses around that mean, given by temperature. In effect, you construct an algorithm that translates model parameters to predicted measurements. Any algorithm can be described as a function. Your idea of how models are concluded by testing causal assumptions is cargo-cult science. It is only partially correct and vastly misleading. In particular, such testing isn't always possible even in principle. You have untestable properties and parts of models that are simply chosen in lieu of better alternatives. By restricting yourself to such simple-minded testing, you blind yourself to large parts of reality even. Many interesting topics have no way of completely controlling conditions for example. They are not arbitrarily repeatable either. Astronomy, economics, psychology, biology...the world is bigger than your approach can account for.
- gwern 3y agoA good compressor also needs to compress data from experiments using randomization. Causal data is also data. I don't really know what more needs to be said there.
- mjburgess 3y agoThere is no such thing as "causal data". A causal model is an interpretation of data. Eg., to say "increasingly energetic motion of molecules leads to increasingly hot water" is an interpretation of a very wide class of equations. It posits the existence of molecules (a scientific discovery), water, energy, motion, heat, etc. and it provides a means of creating equations&measures tied to each of these terms. Science is the production of those interpretations. There is no bare "data" which tells you how reality is. Science isn't "magic trick engineering", it's Explanation. "Compressing tables of data" is something they do in the pseudosciences -- as you've seen, none of it is reproducible: "IQ" is just a compression of survey quizzes. Do you really think it exists? Do you think you can just compress survey results and claim to have an explanatory model of the most complex system in the entire universe? (a person, society, and their joint interaction) etc. ML is a temple to pseudoscience, permitted only because the situations it's used in are engineered and low-risk. The whole thing is a dumb trick. You cannot build models of the world from associations in data: that is called superstition.
- gwern 3y agoYou flip a coin to randomize choice of a treatment and record the results. The coin-flips+results is a stream of binary data that can then be compressed well or poorly. A compressor which has built a correct causal model of the effects, whatever those are, will compress better than one which is unable to and can only blindly predict pretreatment results (or worse, predict conditional on the correlations which were just broken by the coin-flip, thereby actually wasting bits to fix its especially erroneous predictions). This is in line with the compression paradigm. Where do you disagree? Do you think that causal models are completely useless for shortening predictions? Or do you think causality just doesn't exist? > "Compressing tables of data" is something they do in the pseudosciences -- as you've seen, none of it is reproducible: "IQ" is just a compression of survey quizzes. Do you really think it exists? That's not even close to correct about IQ. You can measure it from lots of things which are not 'survey quizzes'; fMRIs, for example.
- guy98238710 3y agoIsn't noise in the data going to dominate output size of lossless compression? Wouldn't linguistics and vision be better off with direct measurements of predictive strength?
- d_burfoot 3y ago(author) Noise certainly affects the compression rate. But you are not concerned with the absolute compression rate, you are only concerned with the relative rate achieved by two theories A and B. Both theories will be negatively impacted to the same degree by the noise, so the comparison still works to select which theory is better.
- harperlee 3y ago…to the extent that the difference between theories dominates over the possible noise variability.
- gwern 3y agoYou can easily quantify the variance and do standard model-comparison/hypothesis-testing if you want statistical-significance levels. For many datasets these days, this is hardly even a consideration: even a 1% compression improvement is clear.
- d_burfoot 3y ago(author here) Recents events in ML make me feel about 2/3 vindicated of the claims made in the book. Based on the book's ideas, I began training LLMs based on large corpora in the early 2010s, well before it was "cool". I figured out that LLMs could scale to giga-parameter complexity without overfitting, and that the concepts developed under this training would be reusable for other tasks (I called this the Reusability Hypothesis, to emphasize that it was deeply non-obvious; other terms like "self-supervision" are more common in the literature). I missed on two related points. Technically, I did not think DNNs would scale up forever; I thought that they would hit some barrier, and the engineers would not be able to debug the problem because of the black-box nature of DNNs. Philosophically, I wanted this work to resemble classical empirical science in that the humans involved should achieve a high degree of knowledge relating to the material. In the case of LLMs, I wanted researchers (including myself) to develop understanding of key concepts in linguistics such as syntax, semantics, morphology, etc. This style of research actually worked! I built a statistical parser without using any labelled training data! And I did learn a ton about syntax by building these models. One nice insight was that the PCFG is a bad formalism for grammar; I wrote about this here: https://ozoraresearch.wordpress.com/2017/03/17/chuckling-a-bit-at-microsoft-and-the-pcfg-formalism/ https://ozoraresearch.wordpress.com/2017/03/17/chuckling-a-b... Obviously, I feel into the "Bitter Lesson" trap described by Rich Sutton. The DNNs can scale up, and can improve up their understanding much faster than a group of human researchers can. One funny memory is that in 2013 I went to CVPR and told a bunch of CV researchers that they should give up on modeling P(L|I) - label given image - and just model P(I) instead - the probability of an image. They weren't too happy to hear that. I'm not sure that approach has yet taken over the CV world, but based on the overwhelming success of GPT in the NLP world, I'm sure it's just a matter of time. In hindsight, I regret the emphasis I placed on the keyword "compression". To me, compression is a nice and rigorous way to compare models, with a built-in Occam's principle. But "compression" means many different things to different people. The important idea is that we're modeling very large unlabelled datasets, using the most natural objective metric in this setting. edit: I used the wrong name in reference to the Bitter Lesson idea, here is the essay: http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html
- gwern 3y ago