6 ms·
Papers with Code
- yukinon 5y agoThis is a great site. It's pretty ML focused which lands a bit outside my interest range, does anyone know of a similar site that has papers from CS as a whole?
- srvmshr 5y agoNot exactly what you need, but this page lists best papers awards (with links) from all major conferences https://jeffhuang.com/best_paper_awards/ https://jeffhuang.com/best_paper_awards/ And here's PapersWeLove Repo with similar sauce https://github.com/papers-we-love/papers-we-love https://github.com/papers-we-love/papers-we-love
- wallflower 5y agoYes. I just resubmitted Papers We Love. Submission activity has dropped off and still a Scrooge McDuck’s embarrassment of riches. https://paperswelove.org/ https://paperswelove.org/
- barefeg 5y agoWhat are your interests? This one has some extra content like company blogs and conferences, though it’s still AI centric https://www.zeta-alpha.com/ https://www.zeta-alpha.com/
- mnks 5y agoGlad you like Papers with Code. Please check [1] for the list of scientific domains we currently support and [2] for CS in particular. [1]: https://portal.paperswithcode.com/ https://portal.paperswithcode.com/ [2]: https://cs.paperswithcode.com/ https://cs.paperswithcode.com/
- rg111 5y agoIt is weird to see Papers with Code on the front page of HN. This site is the bread and butter of each Research Engineers and Scientists working in Deep Learning. You use the site almost everyday. Advanced learners also use the site regularly. You would just think that "everyone knows" and never think of sharing the site on HN.
- fault1 5y agoI think you overestimate how many HN readers are "Research Engineers and Scientists working in Deep Learning".
- p1esk 5y agoHe also overestimates the importance of that site for “Research Engineers and Scientists working in Deep Learning".
- hungryforcodes 5y agoI've never heard of it. Glad to have made its acquaintance.
- visarga 5y agoLet me guess - everyone - means /r/machinelearning and a curated list of people on Twitter?
- criticaltinker 5y agoA couple previous discussions for those interested: https://news.ycombinator.com/item?id=19054501 https://news.ycombinator.com/item?id=19054501 (Feb 1, 2019) 411 points, 23 comments https://news.ycombinator.com/item?id=23391934 https://news.ycombinator.com/item?id=23391934 (June 2, 2020) 304 points, 21 comments
- barefeg 5y agoI’m curious, how does it fit your daily workflow as an engineer? Is it somewhere where you get the “news” for the day? Or do you use it for getting information relevant to your current work projects?
- igorkraw 5y agoYou use it to find the code and data of a paper - since it also lists other implementations - to run additional baselines on Imagenet in order to appease reviewer #3( who has no idea why your paper on convex optimization has nothing to do with this but it's easier to run them than argue with them). Pre-parenthesis part is dead serious, parenthesis part is slightly hyperbolic due to accumulated trauma with bad reviewers
- criticaltinker 5y agoTransformer based architectures and unsupervised pre-training are achieving state of the art results across multiple modalities including NLP, CV, speech recognition, genomics, physics etc - so here's my must read list of recent papers on the topics (along with some of my notes). Happy holidays! [1] Attention Is All You Need (2017) https://paperswithcode.com/paper/attention-is-all-you-need https://paperswithcode.com/paper/attention-is-all-you-need Introduced the Transformer architecture and applied it to NLP tasks. [2] The Annotated Transformer (2018) https://nlp.seas.harvard.edu/2018/04/03/attention.html https://nlp.seas.harvard.edu/2018/04/03/attention.html An “annotated” version of [1] in the form of a line-by-line Pytorch implementation. Super helpful for learning how to implement Transformers in practice! [3] BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2018) https://paperswithcode.com/paper/bert-pre-training-of-deep-bidirectional https://paperswithcode.com/paper/bert-pre-training-of-deep-b... One of the most highly cited papers in machine learning! Proposed an unsupervised pre-training objective called masked language modeling; learned bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers. Bonus: https://nlp.stanford.edu/seminar/details/jdevlin.pdf https://nlp.stanford.edu/seminar/details/jdevlin.pdf See the above slideshow from the primary author, noting the remarkably prescient conclusion: "With [unsupervised] pre-training, bigger == better, without clear limits (so far)" [4] Conformer: Convolution-augmented Transformer for Speech Recognition (2020) https://paperswithcode.com/paper/conformer-convolution-augmented-transformer https://paperswithcode.com/paper/conformer-convolution-augme... Proposed an architecture combining aspects of CNNs and Transformers; performed data augmentation in frequency domain (spectral augmentation). [5] Scaling Laws for Neural Language Models (2020) https://paperswithcode.com/paper/scaling-laws-for-neural-language-models https://paperswithcode.com/paper/scaling-laws-for-neural-lan... Arguably one of the most important papers published in the last 5 years! Studies empirical scaling laws for (Transformer) language models; performance scales as a power-law with model size, dataset size, and amount of compute used for training; trends span more than seven orders of magnitude. [6] Language Models are Few-Shot Learners (May 2020, NeurIPS 2020 Best Paper) https://paperswithcode.com/paper/language-models-are-few-shot-learners https://paperswithcode.com/paper/language-models-are-few-sho... Introduced GPT-3, a Tranformer model with 175 billion parameters, 10x more than any previous non-sparse language model. Trained on Azure's AI supercomputer, training costs rumored to be over 12 million USD. Presented evidence that the average person cannot distinguish between real or GPT-3 generated news articles that are ~500 words long. [7] CvT: Introducing Convolutions to Vision Transformers (May 2020) https://paperswithcode.com/paper/cvt-introducing-convolutions-to-vision https://paperswithcode.com/paper/cvt-introducing-convolution... Introduced the Convolutional vision Transformer (CvT) which has alternating layers of convolution and attention; used supervised pre-training on ImageNet-22k. [8] Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition (Oct 2020) https://paperswithcode.com/paper/pushing-the-limits-of-semi-supervised https://paperswithcode.com/paper/pushing-the-limits-of-semi-... Scaled up the Conformer architecture to 1B parameters; used both unsupervised pre-training and iterative self-training. Observed through ablative analysis that unsupervised pre-training is the key to enabling growth in model size to transfer to model performance. [9] Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity (Jan 2021) https://paperswithcode.com/paper/switch-transformers-scaling-to-trillion https://paperswithcode.com/paper/switch-transformers-scaling... Introduced the Switch Transformer architecture, a sparse Mixture of Experts model advancing the scale of language models by pre-training up to 1 trillion parameter models. The sparsely-activated model has an outrageous number of parameters, but a constant computational cost. 1T parameter model was distilled (shrunk) by 99% while retaining 30% of the performance benefit of the larger model. Findings were consistent with [5]. [10] ProtTrans: Towards Cracking the Language of Life's Code Through Self-Supervised Deep Learning and High Performance Computing (August 2021) https://paperswithcode.com/paper/prottrans-towards-cracking-the-language-of https://paperswithcode.com/paper/prottrans-towards-cracking-... Applied Transformer based NLP models to classify & predict properties of protein structure for a given amino acid sequence, using supercomputers at Oak Ridge National Laboratory. Proved that unsupervised pre-training captured useful features; used learned representation as input to small CNN/FNN models, yielding results challenging state of the art methods, notably without using multiple sequence alignment (MSA) and evolutionary information (EI) as input. Highlighted a remarkable trend across an immense diversity of protein LMs and corpus: performance on downstream supervised tasks increased with the number of samples presented during unsupervised pre-training. [11] CoAtNet: Marrying Convolution and Attention for All Data Sizes (December 2021) https://paperswithcode.com/paper/coatnet-marrying-convolution-and-attention https://paperswithcode.com/paper/coatnet-marrying-convolutio... Current state of the art Top-1 Accuracy on ImageNet.
- 2bitencryption 5y agoQuestion - I get that your run-of-the-mill paper saying "Here we present a novel algorithm for xyz" will usually have the algorithm defined in simple psuedo-code, maybe with an implementation in a "real" language as a proof of concept. But for the many papers describing novel ML models, how does that work? They seem to use images that diagram out the different layers of the model. But is that truly "universal" the way that a psuedo-code algorithm is universal? As in, if the authors use PyTorch (or whatever), can I take the exact model they describe in their paper and apply it in MyFavoriteMLToolkit and achieve similar results? I guess my question is, what are the "primitives" of papers describing ML models? Is saying "convolutional layer" enough, or do they also describe the dozens of hyper-parameters, etc?
- zacmps 5y agoIt depends. Usually a paper doesn't have enough room to mention all of the possible choices in preprocessing, architecture, optimiser, etc. You can usually get pretty close with details just in the paper, but it's not always possible. That's why a large number of journals now have requirements for publishing code and/or pretrained models (if applicable). An annoying trend I've noticed in a number of SotA ML papers in video classification present multiple models and only publish the exact architecture & weights for the smaller models which are only as-good-as SotA (see tiny video networks, X3D for examples).
- liquidmetal 5y agoIn my experience there are many lesser significant hyperparameters that can impact performance when going from the released code to your personal favorite framework. Nothing you can't figure out by reading source code of the two frameworks or by reading the documentation closely. Generally, people don't seem to care about reproducing exact metrics - as long as it is close enough they're happy. You need to dig a bit deeper if you want the full quality.
- criticaltinker 5y agoIt's a good question which might yield a very complex answer depending on how far down the rabbit hole of reproducible science/computation/machine learning you're willing to go. To keep things simple, I'd say the true "primitives" of ML models can be reduced to mathematical formulas. For example, a plain old feed forward network is implemented as matrix multiplication. Sprinkle in a bit of calculus to analytically derive the formula for back-propagating errors (aka training), and you have the basic building blocks of modern deep learning. Convolutions, Transformers, etc are just a bit fancier spins on the same mathematical foundations. Hyper-parameters are essentially tunable variables in a formula. I'd say your instinct is spot on - they are absolutely necessary to capture for reproducible results. If you have the code and the data the answer should be yes. You should be able to take that PyTorch code and translate it to MyFavoriteMLToolkit to obtain numerically identical results. In practice, we face the same universal difficulties as other computer science based research: fighting inconsistencies in software, hardware, all the way down to the physics of the universe with cosmic ray induced bit flips, etc.
- motiejus 5y agoI implemented Wang–Müller algorithm, described it, and embedded the code to the pdf, along with tooling how to generate the example diagrams of the paper (and the whole paper). Everything is in the pdf[1]. Arxiv.org won't accept a pdf with attachments though, so only a stripped-down version will come there (once/if I get an endorsement, fingers crossed). I copied this concept from Joe Armstrong, where he suggested to distribute Erlang modules as PDFs with code files (*.erl) as attachments. "Documentation comes first, and the distribution should prioritize humans". [1]: See Section A.1 of https://github.com/motiejus/wm/blob/main/mj-msc-full.pdf https://github.com/motiejus/wm/blob/main/mj-msc-full.pdf
- p1esk 5y agoI don’t get it, why not just include a link to github in your pdf?
- motiejus 5y agoLinks to external sites have a tendency to rot. I've stumbled upon a number of scientific papers from 2000s that include links to sourceforge for code listings. Most of those are dead now. Github will not be there for ever.
- productceo 5y agoAmazing website. As some noted in comments, widely used in AI research community, but I expect this website will be useful to the broader developer community as well!