Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rvarma
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
Understanding Python's GIL
(rohanvarma.me)
1 points
by
rvarma
7y ago
|
0 comments
2.
▲
Analyzing Stripe's Chargeback Protection Service
(rohanvarma.me)
2 points
by
rvarma
7y ago
|
0 comments
3.
▲
Hessians: A tool for neural network optimization
(rohanvarma.me)
2 points
by
rvarma
8y ago
|
0 comments
4.
▲
by
rvarma
9y ago
Ah yeah, my bad, I should've instead shown the activations after the first layer, since "activation 0" is just the distribution of the random data I started with
5.
▲
by
rvarma
9y ago
I actually think the idea of using leaky ReLUs is interesting, because it'll still provide a small gradient when x < 0, which perhaps may slightly alleviate the vanishing gradients issue
6.
▲
by
rvarma
9y ago
Thanks for your comment! Regarding point 1, I stored both the activations before batchnorm and after since I needed them during the backwards pass. i.e. I stored h_out before and after these operations: h_out = (h_out - np.mean(h_out, axis
7.
▲
Batch Normalization for deep networks
(rohanvarma.me)
34 points
by
rvarma
9y ago
|
12 comments
8.
▲
A comparison of loss functions for machine learning
(rohanvarma.me)
4 points
by
rvarma
9y ago
|
0 comments
9.
▲
Nano Review – A microblog for reviewing academic papers
(nanoreview.xyz)
4 points
by
rvarma
9y ago
|
1 comments
10.
▲
by
rvarma
9y ago
Thanks for the link! It actually provides a really clear and intuitive explanation of the notion of similarity.
11.
▲
Language Models, Word2Vec, and Efficient Softmax Approximations
(rohanvarma.me)
117 points
by
rvarma
9y ago
|
33 comments
12.
▲
Implementing a Neural Network in Python
(rohanvarma.me)
2 points
by
rvarma
9y ago
|
0 comments