5 ms·
Personal sad story, but hopefully relevant: during my recent PhD I worked on a problem where I used a Dirichlet Process in my solution. That paper has been boun
by abhgh 2y ago
Personal sad story, but hopefully relevant: during my recent PhD I worked on a problem where I used a Dirichlet Process in my solution. That paper has been bouncing around for the past few years getting rejected from every venue I have submitted it to. My interpretation is that most reviewers (there are exceptions - too few to impact the final voting) don't understand any non-DL theory anymore and are not willing to read up for the sake of a fair review. This is based on their comments, where we have been told that our solution is complex (maybe? - but no one suggests an alternative), exposition is not clear (we have rewritten the paper a few times - we rewrite it based on comments from venue i to submit to venue i+1 - its a wild goose chase), and in one case, someone said the paper is derivative because it uses Blackwell-MacQueen sampling; their evidence? - they skimmed through a paper we had cited that also used the sampling algorithm. This is like saying a paper is derivative because it uses SGD.
I am on the review panel of some conferences too and it is not uncommon to be assigned a paper outside of my comfort zone. That doesn't mean I cut and bail. You set aside time, read up on the area, ask authors questions, and judge accordingly. Unfortunately this doesn't happen most of the time - people seem to be in a rush to finish their review no matter the quality. At this point, we just mechanically keep resubmitting the paper every once a while.
Sorry, end of rant :)
- aspenmayer 2y agoIs a preprint of your paper available? I looked at your blog a bit and was able to find this, which may be it? > Learning Interpretable Models Using Uncertainty Oracles https://arxiv.org/abs/1906.06852 https://arxiv.org/abs/1906.06852 https://doi.org/10.48550/arXiv.1906.06852 https://doi.org/10.48550/arXiv.1906.06852
- abhgh 2y agoYes, that's the one: https://arxiv.org/pdf/1906.06852 https://arxiv.org/pdf/1906.06852
- aspenmayer 2y agoI copied the DOI for convenience but they’re the same paper. I have no formal math background really so I can’t speak to your methods but I appreciate that you have shared your work freely. Did you have any issues defending your thesis due to the issues you described above related to publishing? Noticed a typo in your abstract: “Maybe” should be “may be” in sentence below (italics): > We show that this technique addresses the above challenges: (a) it arrests the reduction in accuracy that comes from shrinking a model (in some cases we observe ~ 100% improvement over baselines), and also, (b) that this maybe applied with no change across model families with different notions of size; results are shown for Decision Trees, Linear Probability models and Gradient Boosted Models.
- abhgh 2y agoYes, it did come up during my defense, but it was deemed not to be a concern since I had one prior paper [1] (the original one in this thread of work, the paper I linked above was an improvement over it), and my advisor (co-author on both papers) vouched for the quality of the work. Thank you for pointing out the typo - will fix it! [1] https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2020.00003/pdf https://www.frontiersin.org/journals/artificial-intelligence...
- godelski 2y ago> Sorry, end of rant :) Don't be. Issues like this are the reason I haven't defended yet. The fact that an AC didn't laugh at that "critique" is itself indicative of a problem. They're as checked out as the reviewers. I was doing work in a more mathy area and could not get assigned reviewers that understood what was being done. To try to get something through I tried a more popular domain, won a bet with my advisor that I could get SOTA on a very popular dataset in a few months, but I have no compute left. I can beat big labs on one dataset with far less compute, but how can I compete when reviewers want dozens? Even if others weren't held to that standard... There's not enough compute for that. You can always have "more experiments" For review, I set aside hours for each paper, and more the further out of my domain that they are (I'm also very happy to increase my score with a rebuttal and mark lower confidence (I frequently write what would change my mind to help authors). My best post rebuttal ever was "The authors answered all my questions, but due to the lack of novelty I'm lowering my score"). I'll keep doing this, but to be honest, after I defend I have no intention to push to conferences or journals. I just fail to see the value. It has caused me to spend more time rewriting and taking away from research. It just makes me upset and isn't making me a better researcher. I crave for someone to actually _criticize_ my work. I have a good citation count and h-index. My best paper is "unpublished", has hundreds of citations, resulted in a very popular blog post, and years later people are still messaging me for using it in their work. I don't think I'm a top researcher, but I don't think I'm well below the pack. I just hate that my research directions are pigeonholed. That you need to do topics that people care about. That you need to evaluate with large scale. As if we can't have conclusions beyond the empirical. As if this isn't about communicating our work. That I need to write to those that are not "peers" (niche domain experts, as opposed to broader domain experts). As if experiments aren't proxies, but are demonstrations of a product. I think this significantly slows down the progress to AGI since it causes us to railroad to build from large models from big companies, and there is so little interest in anything else. How can we explore more architectures, learning methods, and all that if we're required to get SOTA out of the gate? I don't want to say too much about my work since it is still bouncing around in review and I don't want to dox myself. But I'll say something about a work that I __reviewed__. It was for Neural PDEs. Review was for a workshop, and it was clear to me that this paper was rejected from the main conference. What was not clear is why. Until I got to see the reviews form my peers. Their complaints had the standard "novelty" and "not well written" (it was very well written btw), but the kicker for them was that the datasets were synthetic... Like... what?! Why does that even matter? They're solving equations! Luckily they had low confidence and I got the paper through. I wasn't surprised when a few months later I stumbled upon the paper again and found out it was from Welling's group. > At this point, we just mechanically keep resubmitting the paper every once a while. I really wonder how long it will take conference organizers to recognize that the noise in the review process is a significant contributor to the increasing number of submissions. This seems a rather obvious connection but I rarely hear it discussed. Not to mention that it can damage papers quality (this certainly happened to mine, and I suspect yours). Reviews can improve the papers if the review contains actual critiques. But hey, why do work when no one questions a reject? I feel like mine was more ranty lol. But it helps to not feel alone.
- somethingsome 2y agoJust a note > exposition is not clear (we have rewritten the paper a few times - we rewrite it based on comments from venue i to submit to venue i+1 - its a wild goose chase) Does not mean that the paper is invalid, but maybe the storyline is difficult to follow, the results not easy to interpret, or overall badly written or missing justifications. Even if you take into account the reviews to rewrite it, it doesn't mean the paper is clear and easy to understand. As you noted, researchers need to read material outside of their confort zone, and the publications have shifted in focus. Before you could expect a reader to be familiar to the topic, now you need to educate him as clearly as possible. I picked a random text inside the paper > The workings of the technique itself are presented at a high-level in Figure 2. Annoying to read. > Instead of learning the training distribution directly, which might be expensive because of the dimensionality of the data, we first project the data down to one dimension. Why is that good enough? Justification missing > This is done just once, and is shown in the left panel in Figure 2. Since we are solving for classification, we pick this dimension to be a numeric indicator of how close an instance is to a class boundary. Why is it a good indicator, justification > As a convenient proxy, we train a separate highly accurate probabilistic Ok, references on previous research that show it can work? So in essence, I don't say you need to explain everything, but the text could be more clear on the choices and why they make sense. My gut feeling is that you know and understand what you are doing, but you miss too many justifications that proves your work valuable. I didn't read the whole thing, so maybe I'm missing the picture, but from random sampling on the text I expect the rest to follow the same. While I read the introduction, I don't want to read 'we did that and that and that'. But 'there was this issue, we solve it in this way because this reason ' And following issues->solution->why should give me enough understanding of what you are trying to achieve. Follow-up sections should refine the solutions
- abhgh 2y agoThank you for these comments. I appreciate them and I'll consider them in my next draft. However, I would like to point out a few things; just so that we have the larger picture in mind. Again, I do appreciate you took the time to look up the paper. 1. When I said we revise the paper between two submissions, I wasn't implying it was becoming "better". The message was that there is no general consensus around what should be expanded and what might be concise. Someone believes you should discuss prior work more, someone thinks the main algorithm requires more elaboration, someone wants you to talk more about BayesOpt etc., but you just have <10 pages in the main paper, and putting this stuff in the Appendix, or citing source, doesn't seem to be good enough in many cases (another comment in a sibling thread gives an example wrt GANs, and my experiences have been no different). 2. You say you randomly picked a few sentences to read; that's good for a casual discussion but that should not be how a review process functions. Some of the best reviewers I've encountered (and I hope I am continuing in that tradition) come back to say something like "I see what you're getting at, but your intro. doesn't sell it well enough; think about writing it like this ...". Rejecting based on random skimming is exactly one of the things I'm calling out. Let's face it - like a lot of things, high quality reviewing is hard. It isn't supposed to be quick or easy. 3. Predicting how much to elaborate: this is probably an extension of the first point, but I feel like this has become way harder in the recent years. The rule that mostly works seems to be that if its not a trending topic explain it as much as you can, because cited background material is overlooked. This is unfair for areas that are not trending - the goal of research should be to situate itself closer to "explore" on the "explore-exploit" spectrum, but the review system today heavily favors "exploit". And like I mentioned, a page limit means that the publication game stacked against people not working on mainstream ideas. This should not be the case.
- iamcreasy 2y agoThank you for writing the blog post on Jensen's inequality. It is one of the best introductory material on this topic I've seen.
- abhgh 2y agoThank you!
- joshjob42 2y agoAh, Dirichlet Processes, such lovely things. Reading this paper, I was struck by how obvious most of the solutions were given my own background from grad school benchmarking quantum annealers and other classical solvers for spin lattices (mostly thermal sampling inspired approaches). I'd argue one could do an even better job than the analysis in Anthropic's paper, but it's astonishing how basic questions like "well how sure are we this is real?" just aren't asked seemingly in ML papers. I developed a passion for Bayesian statistics approaches in grad school, and had a lovely time specifically thinking quite a bit about DPs, Bayesian bootstraps, etc. I'm sorry your paper is bouncing around. I think folks underestimate these days the value of really thinking about what you know and how you know it, and how to really model uncertainty, and definitely underrate non-DL approaches to problems.
- abhgh 2y agoThanks, yes, lot of good ideas in ML seem to be slowly vanishing from the collective awareness. I have nothing against the current spate of methodologies which are empirically great - and if one needs proof, I am a "happy customer" at my day job which is mostly DL and a lot of LLMs - but it seems we are buying into a world where it is one versus the other. And this it need not be. Great ideas are great ideas irrespective of age and there is value in preserving them. Anyway, since this thread surprisingly evoked a mini-discussion on Dirichlet Processes (DP), if someone needs an intro, I have tried to balance math and intuition in a description in my thesis: Section 2.2 in [1]. [1] https://drive.google.com/file/d/1zf_MIWyLY7nxEr5UioUQ7KhOQ1_clYYl/view?usp=drive_link https://drive.google.com/file/d/1zf_MIWyLY7nxEr5UioUQ7KhOQ1_... EDIT: I looked at the description and I confess it still has a lot of math (since it is part of thesis). I will probably translate this to be more friendly and put it on my blog.