Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
vmarkovtsev
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
To Build or to Empower: Software Engineering Career Paths Explained
(athenian.com)
2 points
by
vmarkovtsev
4y ago
|
0 comments
2.
▲
Awesome Machine Learning on Source Code
(github.com)
3 points
by
vmarkovtsev
9y ago
|
0 comments
3.
▲
by
vmarkovtsev
10y ago
Me: https://twitter.com/tmarkhor/status/842347955390685184
4.
▲
Multi-GPU Yinyang K-means and K-nn with native R and Python bindings
(github.com)
3 points
by
vmarkovtsev
10y ago
|
0 comments
5.
▲
by
vmarkovtsev
10y ago
Agreed. Sent you a job proposal. We are http://sourced.tech
6.
▲
GitHub contributions graph: 6 handshakes theory and PageRank
(blog.sourced.tech)
5 points
by
vmarkovtsev
10y ago
|
0 comments
7.
▲
Dataset: 452M commits on GitHub
(data.world)
2 points
by
vmarkovtsev
10y ago
|
0 comments
8.
▲
Hercules and his labours – Git repositories burndown
(github.com)
1 points
by
vmarkovtsev
10y ago
|
0 comments
9.
▲
Weighted MinHash on GPU helps to find duplicate GitHub repositories
(blog.sourced.tech)
2 points
by
vmarkovtsev
10y ago
|
0 comments
10.
▲
Hands on the most starred GitHub repositories
(blog.sourced.tech)
1 points
by
vmarkovtsev
10y ago
|
0 comments
11.
▲
Reading PySpark pickles locally
(blog.sourced.tech)
1 points
by
vmarkovtsev
10y ago
|
0 comments
12.
▲
Adding LZO support to Dataproc
(blog.sourced.tech)
2 points
by
vmarkovtsev
10y ago
|
0 comments
13.
▲
What's common between Chef and cooking; WinApi and pokemons?
(blog.sourced.tech)
1 points
by
vmarkovtsev
10y ago
|
0 comments
14.
▲
Jupyter VFS backend for Google Cloud Storage
(github.com)
1 points
by
vmarkovtsev
10y ago
|
0 comments
15.
▲
by
vmarkovtsev
10y ago
I am working on the multi-gpu branch at the moment, but the memory constraints will remain the same (optimizing for the speed at this time). I do see the way to distribute memory across cards though, and it will be the next step. So yes, if
16.
▲
by
vmarkovtsev
10y ago
1. The number of samples must not exceed UINT32_MAX, that is, 4*10^9. The number of clusters must not exceed UINT32_MAX too. Number of dimensions must not exceed 12288 (GPU shared memory constraint). We successfully tested with 4M samples a
17.
▲
by
vmarkovtsev
10y ago
97x improvement is actually very suspicious. Thanks for the article!
18.
▲
by
vmarkovtsev
10y ago
Thanks. Actually, I like DBSCAN a lot and use it often, though I am not much familiar with it's internals. It looks like it is iterative and thus does not fit very well to a GPU. The only way I see is to pick several seed points at sta
19.
▲
by
vmarkovtsev
10y ago
This is an exciting opportunity, thanks, we will study it.
20.
▲
by
vmarkovtsev
10y ago
Here. Please don't bite :)
21.
▲
by
vmarkovtsev
10y ago
Data source is GitHub of course - it's in the article title, so I don't quite get you... We used go-git to fetch every repo out there.
22.
▲
by
vmarkovtsev
10y ago
The age is April 2016. The source is github.com/src-d/go-git-ing
23.
▲
Veles: Distributed platform for rapid deep learning application development
(velesnet.ml)
154 points
by
vmarkovtsev
11y ago
|
14 comments
24.
▲
by
vmarkovtsev
11y ago
This is very similar to what we did in Samsung several years ago: https://velesnet.ml