Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Saurabh_29
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Fast approximate knn using random projections
(github.com)
3 points
by
Saurabh_29
5y ago
|
1 comments
2.
▲
by
Saurabh_29
5y ago
Hi, C++ implementation of approximate nearest knn which doesn't suffer from problem of curse of dimentionality whilst giving theoretical bounds on approximation.
3.
▲
Text to Speech CUDA Programming
(github.com)
3 points
by
Saurabh_29
7y ago
|
0 comments
4.
▲
by
Saurabh_29
7y ago
+1, it is the biggest bottleneck at times
5.
▲
by
Saurabh_29
7y ago
Also, in some cases like small RNN/LSTMs, CPU's can be faster.
6.
▲
by
Saurabh_29
7y ago
I will agree with you. It is a combination of both. Plus I would like to add some extra points, pytorch/tensorflow essentially use the same CUDA/Cudnn libraries, the thing they are developed with the motive to catering wide corner
7.
▲
by
Saurabh_29
7y ago
The main bottleneck is the time spent in adding bias after Conv/dense. A well-optimized code can remove that bottleneck. I haven't done it in my code as it makes it less readable.
8.
▲
by
Saurabh_29
7y ago
The main bottleneck is the data transfer speed between the GPU and the SMs. Also, using tensor core doesn't necessarily apply using half-precision as now NVIDIA supports single-precision operation in Tensorcore too.
9.
▲
Waveglow Inference in CUDA C++
(github.com)
72 points
by
Saurabh_29
7y ago
|
10 comments