Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tasdfqwer0897
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
1.
▲
by
tasdfqwer0897
3y ago
Hey I work at Adept and helped make this! Happy to answer questions. The thing I think is especially neat/notable is how simple you can make the model architecture while still getting good performance. I expect we'll continue to
2.
▲
by
tasdfqwer0897
4y ago
Yeah this is a good point! We are spending a lot of time thinking about reliability and it's true that existing models fall a little flat here. I think ultimately the key to making this work really well is some combination of a) colle
3.
▲
by
tasdfqwer0897
4y ago
Yeah, we did have to custom-build our own benchmarks. And we are not building a chatbot, we're building something collaborative that you can work with to accomplish the stuff you want to do!
4.
▲
by
tasdfqwer0897
4y ago
Thanks - glad you like it! I probably won't get to all of these but let me try a couple: 1. There's a spectrum (sort of) between using full on RL techniques and just doing sequence modeling. We're trying to pick a reasonable
5.
▲
by
tasdfqwer0897
4y ago
We used a combination of human demonstrations and feedback data! You need custom software both to record the demonstrations and to represent the state of the Tool in a model-consumable way.
6.
▲
by
tasdfqwer0897
4y ago
Yes! We plan on putting out a more detailed technical post soon.
7.
▲
by
tasdfqwer0897
4y ago
Hey, I helped make this! Happy to answer any questions.
8.
▲
Neural Networks Can Use Software Tools in Response to User Commands
(vimeo.com)
1 points
by
tasdfqwer0897
4y ago
|
0 comments
9.
▲
Video of a Neural Network Using Software Tools
(twitter.com)
1 points
by
tasdfqwer0897
4y ago
|
0 comments
10.
▲
Adept AI Labs
(twitter.com)
1 points
by
tasdfqwer0897
4y ago
|
0 comments
11.
▲
by
tasdfqwer0897
5y ago
> do you see improvements in Transformer or attention based architectures as essential... I do personally, but there is some disagreement about this in the field. In fact, I would go further and say that (in addition to using large pre-
12.
▲
by
tasdfqwer0897
5y ago
I personally agree that this experiment is evidence that there are certain problems that cannot be solved simply by making the models bigger, and one of the main research questions I'm interested in is what we need to do to elicit more
13.
▲
by
tasdfqwer0897
5y ago
Unfortunately not, but we do release both the programming dataset and the math questions dataset, so in principle you could try those out with one of the open-source models from e.g. huggingFace.
14.
▲
by
tasdfqwer0897
5y ago
I think it might be a mistake to think that the model is not confident because its response is something a human might say if they were not confident. The model is 'just' completing the prefix text with something that has high lik
15.
▲
by
tasdfqwer0897
5y ago
Hey, I am one of the lead authors of this paper. Happy to answer questions. This is a twitter thread going over the main results: https://twitter.com/gstsdn/status/1427794393373626368
16.
▲
Program Synthesis with Large Language Models
(twitter.com)
3 points
by
tasdfqwer0897
5y ago
|
0 comments
17.
▲
by
tasdfqwer0897
6y ago
We have been working on this recently on the Google Brain team. We are working both on synthesizing programs from scratch (see https://arxiv.org/abs/2002.09030 for example) and on understanding computer programs using
18.
▲
by
tasdfqwer0897
7y ago
This actually might have interesting connections to ideas from differential privacy. Maybe the work is derivative of a particular training image if we can easily predict the presence or absence of that training image given only the trained
19.
▲
by
tasdfqwer0897
7y ago
So if you wave your hands enough, it seems like maybe there's an argument to be made that the weights of a trained GAN somehow correspond to a 'compilation' of the training data as it's defined in this doc: https:/
20.
▲
by
tasdfqwer0897
7y ago
So you are worried that your existing attribution method is too focused on 'obvious' attributes and you want to see if you can make it focus on less obvious things? IIUC, that's something that's been looked at in the ML
21.
▲
by
tasdfqwer0897
7y ago
Someone on the machine learning reddit asked me this: > Question: How does copyright work for GAN output? If I input 300,000 copyright protected photos of celebrities and generate images of new celebrities that do not exist, are the gene
22.
▲
by
tasdfqwer0897
7y ago
Hmm, I'm not sure what you mean by applicable loss functions? I'll answer what I think you're asking and you can tell me if I got it wrong: There's been a lot of effort spent on coming up with different loss functions fo
23.
▲
by
tasdfqwer0897
7y ago
Hey, I wrote this! Happy to answer questions.