5 ms·
So, how many bits are you giving the deep learning machine? If you give it enough bits, you aren't predicting anything -- you're just feeding it experimental da
by sidek 10y ago
So, how many bits are you giving the deep learning machine? If you give it enough bits, you aren't predicting anything -- you're just feeding it experimental data which it more or less spits back out.
For an elegant solution, one would need to have deep learning that also optimises on a constrained set of bits.
I feel like in this scenario, humans perform much better. The main strength of computer learning is being able to harness a massive dataset and really good storage capabilities.
Also, you would need to impose a "consistency" constraint on the computer which would be hard to do. Like a computer might say if (mass > 5) do this; else do that. And that is valid computationally. We do that in our split between gr/qm. But in some sense we feel this is wrong physically. The universe shouldn't run on arbitrary if statements.
So I think the answer is just no: the computer algorithms we have today can't handle the problem's constraints.
- visarga 10y agoEnough bits to make it efficient at predicting. When it predicts better than the current theory, it starts getting closer to having enough bits. That's the beauty of ML, you don't need to worry about these details if it gives good accuracy. My intuition was that there could be different ways to explain the laws of Physics that don't look like the current ones which evolved based on human intuition, math and language ability. A non-anthropocentric Physics if you will.
- sgt101 10y ago>That's the beauty of ML, you don't need to worry about these details if it gives good accuracy. I get very alarmed by this. At work we have several examples of ML systems that have done good things for many years before suddenly and inexplicably blowing up and producing nonsense. Our folk explanation (as we have failed to produce anything resembling a proper one) is that the models that are captured in some cases appear to replicate reality but are deficient of some fundamental part of it which later comes into play destroying there predictive power. The domain theory changes in a sense, in another sense the driver was there all along but just hadn't featured in this part of the regime.
- ggggtez 10y agoIn ML, they might say you were overfitting. Predicting is all well and good, but predicting too well can indicate the machine hasn't really learned anything. It's just spitting back nearly identical information as the original. It sounds like that is what this black box is supposed to do. All you need for a perfect black box is every possible data point...
- sgt101 10y agoSo we have instance based learning which basically is about approximating this, and support vector machines which abstract the idea into representing the space of decisions represented by all the data points via a kernel which is a mapping function. The problem I see which deep learning seems to be "winning on" is that the space of instances that you have does not either abstract a meaningful theory of the distribution of these points (therefor predicting outside of this space) or describe things exterior to that space - but that exist, or might exist. A scientific theory is considered valid if it predicts things that have not been seen yet (and you can then test that, hence the post today about "string theory : enough"), I think IBL and SVMs can't represent the unseen, I used to work using something called inductive logic programming (I used a tool called Progol that Steven Muggleton made, it was good!) and I thought that that did produce such insights, sometimes, but it turned out that often these were actually me due to fiddling. Some play projects I did recently made me think that conv. nets were doing the same sort of things, but it was /is harder for me to catch and show this (and as I say, I ended up thinking that I'd often fooled myself with ILP, even though it was very useful). The old skool difficulty I have with deep nets is that I used to think in terms of structural risk minimization vs empirical risk minimisation when dealing with overfitting. The idea was that if you had the right size of information store in your learned system and it was experimentally producing results that showed it was an effective predictor you could say that it was generalizing properly. Deep nets seem to me to have all the information storing capability of the domains they address and I worry.... But I am shocked, shocked, by how well they seem to work.