5 ms·
Since this is hacker news I'm sure there's at least a few people who will read this that are the ML research engineers who write papers on their architectures w
by MathYouF 4y ago
Since this is hacker news I'm sure there's at least a few people who will read this that are the ML research engineers who write papers on their architectures which achieve SoTA results on different benchmarks, or establish new datasets, or publish new tooling like different loss functions/optimizers (which usually are published alongside some new SoTA results achieved by integrating them into a current SoTA architecture).
I've wondered from that crowd (the productive crowd) if they get any value from papers like these, if reading them ever gives any insight. I don't seem to ever glean very much from them, whereas just reading the formerly mentioned papers and thinking for myself and experimenting seems to give me the intuition I personally need.
I'm starting to feel there's a whole secondary field within ML of people who publish 50+ page math-heavy theory only that yields no actual results nor value in achieving those results, but maybe I'm wrong. I'll ask around at the conferences this year. I'd be willing to read papers like these if it turned out it'd help me but as time is limited, I feel I'm likely better off reading more papers that show results and their architectures (which usually have a bit of light speculating/intuiting mixed in as well) than reading this.
Also, no offence meant to the writers of the paper who might read this. If you have any thoughts on this line of inquiry as well I'd be interested to hear them. Maybe there's a lot of practical value I'm missing out on by not taking the time to read and understand papers like these.
- scribu 4y agoThe way I see it, the utility of theory is raising the baseline. It guides you away from dead ends, so you don't waste time on approaches that could never work. It's necessarily high-level, so you still need to learn about specific approaches to get practical things done.
- MathYouF 4y agoI think I'd need an author of one of those (SoTA) papers to walk me through how reading part of a 50+ page theory paper helped them achieve any marginal improvement on their path towards productive contributions for me to really understand or believe it. If I could have even a single example explained to me I think I'd get it.
- mellavora 4y agoMaybe not reading it, but did you consider that writing that paper helped them clarify their thinking on the subject? And based on that clarity, they were able to then go on and make immediately productive contributions?
- MathYouF 4y agoI'm a bit lost on who the 'they' are and what the 'paper' is in this case: The authors of this paper writing this paper, or the authors of SoTA papers writing SoTA papers. A third option you may have meant which makes sense to me in context is the authors of this paper going on to write SoTA papers. I actually looked through the published works of the first two authors and didn't find any practical work, all similar very long form math notation heavy theoretic papers, which supports what I initially worried about that there's basically two diverging branches of ML papers, one of which I'm skeptical about the practical value of. I have actually not seen any examples of first authors on theory papers going on to publish SoTA results on applications, seeing such a thing I would also consider pretty convincing evidence of the utility of these papers and the symbiotic unification of the two groups. I wonder if there's been similar concerns raised historically, like at the advent of electricity, between practical engineers focused on creating groundbreaking applications versus people still focused on theory, and in retrospect what contributions continued to be gained by the theoretical work afterwards in those cases.
- inciampati 4y agoThese syntheses can give powerful grounding to the ideas you get intuitively across dozens or hundreds of applications driven papers. They are useful.
- nabla9 4y agoYes. When you pick up a SoTA algorithm and try to apply it into a particular problem, ... it does not work or does not work well enough. You must modify or fix it. Combine it with other approaches. Be able to know what algorithms to pick. Without knowing how the thing works, what the internal dynamics is, what are the bottlenecks and limitations you get nowhere. Trying to fix problems by blindly tweaking parameters is not going to work. You don't have to be able to analyze and do research on internals by yourself, but you must understand the papers and be able to get the idea. Just like engineer doing signal processing must know lot about Fourier transform, wavelets, Laplace transforms even if they forget the details and forget some equations.
- aaraujo002 4y agoIMHO these books are not used for research. In order to contribute to research you need to be at the cutting edge of the existing contributions which are the most recent papers. On the other hand, these kind of books are useful because there are a huge aggregation of past research and it is based on these kinds of ressources that professors make their courses and that undergrads and graduates students learn.
- axg11 4y agoDirect value: no. Indirect value: there are many facets to every field of maths/science/philosophy and each facet makes progress at a different pace. This type of paper doesn’t have any impact on my day-to-day thinking but I don’t doubt that this approach will eventually lead to better methods and understanding of deep learning. Unfortunately that understanding could take 50 years for all we know. In the meantime, progress in applications of deep learning are outpacing progress in theory by several orders of magnitude. The original GAN paper by Ian Goodfellow was published in 2014. Just 8 years later we have ML models that can generating convincing images from text prompts, arguably beyond what 99.9% of humans are capable of. Theory has no chance when engineering is moving so rapidly.
- goodtraveler 4y agoAnd yet none of these systems can do anything beyond generating statistical patterns adapted to our visual sense. It's certainly a worthwhile achievement but I have yet to see anyone make a convincing case that there is a way to scale these systems to anything resembling intelligence. My current test for whether these systems are showing us how intelligence actually works in the human brain is a convincing proof of Cauchy's integral formula and I'm certain this benchmark will remain unsolved for the foreseeable future.
- aesch 4y agoWhy would your benchmark for intelligence in the human brain be something that less than 1% of humans are capable of achieving? To me that type mathematical proof is an example of a very specific type of intelligence rather than general intelligence.
- goodtraveler 4y agoBecause if an algorithm is capable of setting up the conceptual machinery for making sense of and explaining Cauchy's integral formula then that algorithm is revealing something about the structure of the human brain and as such it would be a good benchmark for understanding how human intelligence actually works (not just in the 1% of people). Moreover, mathematics is a fantastic proving ground for testing algorithms that purport to explain symbolic intelligence because if something can not work in a mathematical setting/context then there is no hope it will ever work in the real world since real world intelligence is much more than symbol shuffling. Cauchy's integral formula is a reasonable proxy for proving understanding of symbolic systems and not just juggling their statistical properties/associations. If you think that Cauchy's integral formula is too complicated then there are probably simpler problems that would also serve as reasonable proxies of symbolic understanding, e.g. elementary group theory and linear algebra.