7 ms·
I found the science exams results interesting and skimmed the paper [1]. They report an accuracy of >90% on the questions. What I found puzzling was that they h
by FiberBundle 6y ago
I found the science exams results interesting and skimmed the paper [1]. They report an accuracy of >90% on the questions. What I found puzzling was that they have a section in the experimental results part where they test the robustness of the results using adverserial answer options, more specifically they used some simple heuristic to choose 4 additional answer options from the set of other questions which maximized 'confusion' for the model. This resulted in a drop of more than 40 percentage points in the accuracy of the model. I find this extremely puzzling, what do these models actually learn? Clearly they don't actually learn any scientific principles.
[1] https://arxiv.org/pdf/1909.01958.pdf https://arxiv.org/pdf/1909.01958.pdf
- wrs 6y agoI would be interested in hearing the results from humans presented with adversarial answer options. You may say that a machine learning correlations between words isn’t really learning science, but I wonder how many human students aren’t either, just pretty much learning correlations between words to pass tests...
- FiberBundle 6y agoThey do give an example of a question, in which the model chose an incorrect answer in the adversarial setting: "The condition of the air outdoors at a certain time ofday is known as (A) friction (B) light (C) force (D)weather[correct](Q) joule (R) gradient[selected](S)trench (T) add heat" I assume this might be characteristic for other questions as well, although I don't know anything about the Regents Science Exam and whether there are multiple questions about closely related topics.
- taneq 6y agoThat’s a terribly worded question anyway. Of the original answers, ‘weather’ is the least worst but it’s still vague.
- mannykannot 6y agoIt is a well-worded question for its purpose. The whole point is that, of all the options given, only one is justifiable (and it does not require a tendentious stretch to justify it, either.) Even “light” (which was not chosen) only applies half the time, on average. This is a valid test of natural language understanding.
- rvense 6y agoRemember when IBM went on Jeopardy? There was a question about which Egyptian pharaoh. A human with some knowledge of history might mix up Ramses and Seti, or whatever, or just not know the answer, but know that they didn't know. Watson answered "What are trousers?" with supreme confidence. Jeopardy is fun and games and it was great for the blooper reel, but they're trying to sell this stuff to diagnose cancer and guide police efforts. Failure modes are kind of important.
- jacobwilliamroy 6y agoMost multiple choice math problems can be completely circumvented by simply finding the digital root of the expression in the problem. I was surprised to find this to be true, even on college entrance exams.
- flir 6y agoSeriously? Can you explain the mechanism? 'cos that sounds like - well, numerology, to be honest.
- jacobwilliamroy 6y agoThe digital root of the problem expression will be the digital root of the answer, because the answer IS the problem expression. They're the same number, just with all the pieces jumbled up. Comparing against digital roots of answer choices will usually eliminate most of the wrong answers. Any remaining wrong answers tend to be obviously wrong. 2 + 2 != 13 obviously.
- plutonorm 6y agoCan you give an example? I am intrigued, although suspect you are making a joke!
- deleted 6y ago[deleted]
- jacobwilliamroy 6y agoYou can check the properties of digital roots on wikipedia: https://en.wikipedia.org/wiki/Digital_root#Properties https://en.wikipedia.org/wiki/Digital_root#Properties Example from ASVAB practice math test: (x+4)(x+4) = A. x^2+16x+8 B. x^2+16x+16 C. x^2+8x+16 D. x^2+8x+8 Since we don't have to solve for X in this problem, we can just assume x is 1, which would make the digital root of the problem expression 7. Assuming x = 1, the digital roots of A, B, C and D are 7, 6, 7 and 8 respectively. C is the answer because its last term is the square of 4.
- teej 6y agoThat’s the thing. Machine Learning is a misnomer, the models don’t “learn” anything about the domain they operate in. It’s just statistical inference. A dog can learn to turn left or to turn right for treats. But they don’t understand the concept of “direction”, their brain isn’t wired that way. Machine learning models perform tricks for treats. The tricks they do get more impressive by the day. But don’t be deceived, they aren’t wired to gain knowledge.
- plutonorm 6y agoMy god if I hear this argument one more time I'm going to pop. What on earth gives you the idea that you are something more than statistical inference?
- teej 6y agoYes, it’s widely believed that statistical inference is a part of how the brain operates. But we have barely scratched the surface in our understanding how the human brain works. Do you honestly believe statistical inference completely explains a human’s ability to learn?
- plutonorm 6y agoOf course.
- randcraw 6y agoBecause it is possible to invent or imagine new things without resorting to endless random walks or blind trial and error. Induction isn’t probabilistic. It’s the origin of all discovery and based in sussing out what’s important in patterns, selectively proposing causal mechanisms and testing those hypotheses by following priors and implications — an essential basis for original and creative thought that no purely probabilistic engine can employ.
- FeepingCreature 6y agoYou think a probabilistic engine can't model a human doing induction? At a certain level of capability, mere pattern imitation closes upwards.
- 0-_-0 6y agoAdversiarial training means that they specifically search for answers that the network would misunderstand. If this only leads to a 40 percent loss (the network still answers correctly 50% of the time) I still consider that remarkable. Choosing the best from 8 answers where some of them were adversarially derived should be equivalent to choosing the best from all possible answers, of which there could be tens of thousands. How would a human do in that situation? Although the kinds of mistakes the network makes seem like a mistake you would never do (i.e. you wouldn't call the condition of air outdoors gradient), the opposite could also be true, that it would easily answer questions you would have a problem with.
- mannykannot 6y agoI admire the way that you have reinterpreted the issue, but I don’t buy it. For one thing, I can imagine a knowledgeable human working through ten thousand alternatives and picking the best, (if, objectively, there is one) - it would just take a long time. On the other hand, It does not seem obvious to me that current NLP systems would do better than humans on such a task, and nor is it obvious to me that one can conclude that from the fact that the machine sometimes ignores adversarial examples (if the assumption is that most of the ten thousand would be de-facto adversarial, that is a lot to not make one mistake on; that’s not so much a problem for a human with an understanding of the issue, and who might well come up with the right answer unprompted.) This is probably all moot, however, as the most obvious way to compare the machine results to those of humans is to give both the actual tests in question, rather than substitute a dubious “equivalent” test.
- MiroF 6y agoAdversarial examples probably exist for humans too, they are just less easy to find since we can't easily backprop through human judgements like we can a bunch of array multiplications and dot products. Doesn't mean that they are not "actually" learning any more than we are.
- callmekit 6y agoThese are "adversarial answers" - additional answers to a multi-choice test. If you know the correct answer to the test, adding more possible answers shouldn't make you change your selection.