26 ms·
Hey, I am one of the lead authors of this paper. Happy to answer questions. This is a twitter thread going over the main results: https://twitter.com/gstsdn/s
by tasdfqwer0897 5y ago
Hey, I am one of the lead authors of this paper.
Happy to answer questions.
This is a twitter thread going over the main results:
https://twitter.com/gstsdn/status/1427794393373626368 https://twitter.com/gstsdn/status/1427794393373626368
- ipsum2 5y agoI didn't see it in the paper, but is the code and model going to be released? Would be great to play around with this.
- tasdfqwer0897 5y agoUnfortunately not, but we do release both the programming dataset and the math questions dataset, so in principle you could try those out with one of the open-source models from e.g. huggingFace.
- muds 5y agoGreat work! I found the results in Fig 16 pretty interesting [0]... From response 1, it seems that the model has very little confidence in its decision but get the correct answer while, on the contrary, in response 3, the model seems very confident in its incorrect answer. Is this usually a trend that you see with large models? How hard is it, generally, to make such models "aware" of their own shortcomings? [0]: https://arxiv.org/pdf/2108.07732.pdf#figure.caption.21 https://arxiv.org/pdf/2108.07732.pdf#figure.caption.21
- tasdfqwer0897 5y agoI think it might be a mistake to think that the model is not confident because its response is something a human might say if they were not confident. The model is 'just' completing the prefix text with something that has high likelihood from its perspective, so it may just be used to, for instance, seeing people hedge in similar conversations it has read in its training data. More generally, whether these models are well-calibrated (that is, they know what they don't know) is an important area of research. I don't have references offhand, but I think it's true broadly speaking that these larger pre-trained models do tend to be better calibrated.
- criticaltinker 5y agoFirst, congratulations to you and the other authors - this is a very informative and accessible paper with impressive results. > On both datasets, we find that synthesis performance scales log-linearly with model size This "scaling law" observation is starting to become a major trend in NLP [1] and other other modalities such as speech recognition [2] and protein structure/function prediction [3]. Do you have any insight or commentary to offer regarding the next step in improving program synthesis? For example, will techniques to scale up model size continue to be a primary focus (eg as in [4]), or do you see improvements in Transformer and attention based architectures as essential for pushing the limits of what has been achieved by you and your colleagues? > We find that even our best models are generally unable to predict the output of a program given a specific input. What do you think about leveraging unsupervised training to improve program synthesis? Could synthesized programs be executed on generated input in a way that supports contrastive learning [5]? Thanks in advance for your time and comments here. [1] Scaling Laws for Neural Language Models https://arxiv.org/abs/2001.08361 https://arxiv.org/abs/2001.08361 [2] Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition https://arxiv.org/abs/2010.10504 https://arxiv.org/abs/2010.10504 [3] ProtTrans: Towards Cracking the Language of Life’s Code Through Self-Supervised Learning https://arxiv.org/abs/2007.06225 https://arxiv.org/abs/2007.06225 [4] Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity https://arxiv.org/abs/2101.03961 https://arxiv.org/abs/2101.03961 [5] A Simple Framework for Contrastive Learning of Visual Representations https://arxiv.org/pdf/2108.00587.pdf https://arxiv.org/pdf/2108.00587.pdf
- tasdfqwer0897 5y ago> do you see improvements in Transformer or attention based architectures as essential... I do personally, but there is some disagreement about this in the field. In fact, I would go further and say that (in addition to using large pre-trained models) we will need methods of training that are pretty substantially different in order to elicit robust reasoning behavior. Even supposing I'm wrong about this, if you go and look at the scaling plots in figure 3 and try to figure out how big your model would need to be in order to be solving most of these problems, you'd get a really big number. Even if you had such a big model, it would still require post-processing of the samples to actually get the right answers. From the perspective of applications, that's fine, but it's a little unsatisfying from the perspective of studying intelligence. Even with those caveats (!) these problems aren't that hard compared to general software engineering tasks... > What do you think about leveraging unsupervised training to improve program synthesis? Could synthesized programs be executed on generated input in a way that supports contrastive learning [5]? I think this is an interesting idea and someone should try it! I do think that, even restricting our attention to just getting neural networks to execute programs, that we will need to do something a little more drastic to robustly get the results we want.
- YeGoblynQueenne 5y agoHello Augustus and thanks for posting the link to your co-authored paper on HN. I will have to read the paper more carefully however having quickly scanned the paper it seems that it only reports empirical resuls. In particular, there seem to be no theoretical results about lernability of programs from natural language specifications using large language models. To make it more plain - how do we know these techniques work as well as reported on problems other than the ones in the dataset introduced in your work? Note that I'm not asking why you introduced a new dataset, this seems to be motivated in the abstract. I'm asking: how do we know how well this kind of thing works ("this kind of thing" being what it says in the title) in the general case?
- trainwithsgd 5y agoInteresting read! How would you differentiate your paper from OpenAI's Codex paper (https://arxiv.org/abs/2107.03374 https://arxiv.org/abs/2107.03374)? They also show smooth scaling with parameters, that repeated sampling works, and that BLEU score correlates poorly.