4 ms·
The author is overconfident and underestimates library writers to a fault. Packing a bunch of dot products into a matrix is only the fastest way to compute a ba
by queuebert 5y ago
The author is overconfident and underestimates library writers to a fault. Packing a bunch of dot products into a matrix is only the fastest way to compute a batch of cosine similarities if you are relegated to a parallel matrix product function.
If you write the procedure in raw CUDA, for example, it is faster to simply broadcast the dot products across threads. That is exactly what the matrix multiply is doing except without the overhead of matrix creation and with potentially greater locality of memory accesses.
Edit: I do think it is worthwhile to write your own code so that you better understand what is happening under the hood, so good on the author for that.
- Blackthorn 5y agoThe author is a library writer!
- queuebert 5y agoMy mistake. I took this statement in isolation: "...consider that machine learning libraries are frequently written by grad students on their path to discovery. It's a domain expert with poor programming skills. Or it might be a case of a good programmer who only barely understands the domain…"
- klyrs 5y ago> The author is overconfident and underestimates library writers to a fault. Clicking around the website a bit, it would seem that the author is a library writer (some specifically targeting GPUs), and the article is a plug for the books he's written about numerical analysis. How confident are you?
- queuebert 5y agoNot very confident in general. But I have written a LOT of raw CUDA code.
- inimino 5y agoThis pattern comes up a lot in comments on this type of post. Author makes a dumb but easy-to-follow example to make a point, and then someone points out how the example is dumb. Try to avoid getting hung up on examples -- in this case nobody reading this post should come away thinking they now know more about cosine similarity than a library author. I dare say the author's main points were: - it's not that much code - it's not terrifyingly hard to understand - you might open up performance benefits in your specific circumstance just by knowing how it's calculated after writing the code to do it yourself
- exmadscientist 5y ago> in this case nobody reading this post should come away thinking they now know more about cosine similarity than a library author However, if this post leaves you open to the option that you might know more about $THING than a library author... that is probably a very healthy possibility to consider. (Consider. Some are that grad student who got assigned to write the library because they were too useless to do anything else. Others are David M Gay.)
- inimino 5y agoIndeed! That you should be optimistic about it is the best thing to take away from the post.
- username90 5y agoBeating library implementations is usually very easy since you know your use case perfectly and they don't. You need to be way less competent than them to not be able to do something better when you write a tailor made solution. The main reason not to do it is that it takes time and makes the code harder to get into for new hires.
- queuebert 5y agoLest you think mine was a shallow, knee-jerk analysis, this is my area of subject matter expertise. I get all the nuances, but I just felt like defending library authors, whom he unjustly threw under the bus, because that stuff is really hard. Like really hard. Also, if you are using libraries written by grad students, then you probably need to find better libraries. All of the major battle-hardened ones like BLAS, LAPACK, CUDNN, TF, FFTW, etc., have had major big brains looking at them for years. Decades in some cases.