5 ms·
...out of all the things I expected to see from a Tesla stream, I definitely did not expect a 25 minute discussion on the cost of 32bit additions, dot products,
by 2bitencryption 7y ago
...out of all the things I expected to see from a Tesla stream, I definitely did not expect a 25 minute discussion on the cost of 32bit additions, dot products, sram bandwidth, and chip design, at the level of a third year college hardware course.
- mattrp 7y agoIf you’ve listened to their earnings calls this is par for the course. Say what you will about his tweets, but I find these discussions fantastically informative and transparent.
- isoprophlex 7y agoabsolutely glorious, they're taking questions. first question: can you use anything else besides ReLU activations in your neural net engine? this is wonderful
- akhilcacharya 7y ago...is it? I don't know the context here but it feels like they're trying to ML-splain Karpathy?
- 2bitencryption 7y agowhat a funny question. "you say you can put bananas in your smoothie. just a thought, but perhaps, might I ask, do you have the capability to put strawberries as well, or are we not there yet?"
- pixelpp 7y agohaha
- chrisa 7y agoThat was my first instinct to that question as well, but I think he was just asking if ReLU was required by the hardware design - or if it was possible to use other activation functions as well. If the the ReLU was part of the hardware itself somehow, then it wouldn't be possible to use tanh or sigmoid (which may be better in certain situations); so I think he was just asking if ReLU was required, or if there was flexibility allowed in the activation function.
- 2bitencryption 7y agoah, makes sense. guess I was the fool :P
- cr0sh 7y agoI'm not a NN expert, but based on what I have found, the point of using RELU instead of other activation functions is what is called the problem of "vanishing gradients". Basically (IIRC), during backprop the error difference gets ever smaller the further back in layers you go, ultimately getting "lost in the noise", making learning in the earlier layers more difficult to impossible. I'm not saying RELU is the only option to make this work, or that it's the only activation function that provides a "fix" for the issue; I'm sure there are other ways to deal with vanishing gradients that I don't know about. I also lack the mathematical knowledge as to why RELU helps in this manner, but I suspect something having to do with the lack of "asymptotic structure" approaching the extremes (I don't know what the proper term would be). Or maybe it allows for some form of "forgetting", in the prevention of multiplying very small numbers (such values just go to zero ultimately)? Maybe someone else here with the knowledge can explain it better, and we can both learn...?
- Qub3d 7y agoReLu is useful, but the question as asked was more about other (also essential) functions for other steps in RNNs -- in particular, sigmoid "squash" functions, as well as MaxPooling.
- deleted 7y ago[deleted]
- DeonPenny 7y agoI keep saying this he knows what he doing. You can tell he knows what's happening. Which is rare for a CEO to know this level of detail.
- selectodude 7y agoWhich might explain why he's a crummy CEO. If your job is to strategize for the entire company and you're busy learning the minutiae that you spend a lot of money for other people to worry about, maybe you need to consider using your time a little better. If you want to build neural nets, keep your stock, fire yourself and go build neural nets.
- DeonPenny 7y agoBut then you have the major car company problem and the phone companies before steve jobs. You have people who don't know what's possible, don't know who are the best people to hire, and can guide a vision. if you know capacitive touch screens are possible no one has ever tried it makes it easier to implement. Same with google and how their CEO are all engineers. So pitching the CEO on AI project becomes easy
- outworlder 7y ago> If your job is to strategize for the entire company and you're busy learning the minutiae that you spend a lot of money for other people to worry about, maybe you need to consider using your time a little better This is what everyone that is not Tesla or SpaceX are doing. And have been doing for a long time. If the CEO is not an engineer at heart, what are they? I seriously doubt Elon Musk has more engineering knowledge than the people he hires on their specific fields. However, he can make pretty well informed strategic decisions if he knows WTF the engineers are talking about without taking their word – not even that, as explanations have to be dumbed down. This is not a new thing. Bill Gates was like that (1). Steve Jobs was no dummy and had an engineering background, but not at the same level – he did parter with a genius engineer, however. I think Musk is doing the right thing. > Which might explain why he's a crummy CEO That's quite debatable, I'd say. Isn't he getting results? (1) https://www.joelonsoftware.com/2006/06/16/my-first-billg-review/ https://www.joelonsoftware.com/2006/06/16/my-first-billg-rev...