5 ms·
Paper: https://arxiv.org/abs/1903.03129 https://arxiv.org/abs/1903.03129 Code: https://github.com/keroro824/HashingDeepLearning https://github.com/keroro824/Ha
by exegete 5y ago
Paper: https://arxiv.org/abs/1903.03129 https://arxiv.org/abs/1903.03129
Code: https://github.com/keroro824/HashingDeepLearning https://github.com/keroro824/HashingDeepLearning
Edit: Here is the new paper referred to https://arxiv.org/abs/2103.10891 https://arxiv.org/abs/2103.10891
Here is the new code: https://github.com/RUSH-LAB/SLIDE https://github.com/RUSH-LAB/SLIDE
The paper I linked above was the first paper of which this is supposed to be an improvement.
- guywhocodes 5y agoThis is the initial SLIDE paper, this is supposed to be an improvement on top of SLIDE. Having looked at and run the original implementation I'm a lot less skeptical than the average post here. It was quite fast already and should be considered a toy implemenation.
- djrogers 5y agoWhat do you mean by ‘toy implementation’?
- guywhocodes 5y agoPoC for fun
- gumby 5y agoDid you mean, "should not be considered a toy" or "this original was a toy/PoC and this new paper is much better"?
- guywhocodes 5y ago> this original was a toy/PoC and this new paper is much better This one hundred percent, the original has very few optimizations over a completely naive implementation. It uses MPI and huge pages, that's essentially it.
- foerbert 5y agoI think the reason you're getting so many questions about the 'toy' statement is that often toy programs are simple in large part because they omit handling complex or difficult cases. I don't think I've ever heard the term applied to a complete problem that is merely implemented naively. As a result, saying it's a toy implementation is making people think you mean the speedup is a result of simply handling a simple portion of the problem, rather than thinking you mean even a naive implementation is quite fast.
- d110af5ccf 5y agoIs there anything preventing this algorithm (or a substantially similar one) from being used on the GPU?
- 0-_-0 5y agoAfter a quick read of the paper, no. You could adopt this to the GPU (which would require the hashes work on groups of neurons instead of individuals) and might get a similar speedup. Locality sensitive hashing in fact seems like a primitive attention mechanism, with proper attention implementation you could get maybe even better results.