4 ms·
I quickly read through the paper. One thing to note is that they use the Frobenius norm (at least I suppose this from the index F) for the matrix factorization.
by mo_42 3y ago
I quickly read through the paper. One thing to note is that they use the Frobenius norm (at least I suppose this from the index F) for the matrix factorization. That is for their learning algorithm. Then, they use the cosine-similarity to evaluate. A metric that wasn't used in the algorithm.
This is a long-standing question for me. Theoretically, I should use the CS in my optimization and then also in the evaluation. But I haven't tested this empirically.
For example, there is sperical K-meams that clusters the data on the unit sphere.
- deleted 3y ago[deleted]
- nerdponx 3y agoI think that's kind of the point of the paper. The model is based on un-normalized dot products, and wasn't deliberately designed to produce meaningful cosine similarities. They are showing that, in that case, cosine similarities might be arbitrary and not as useful as people might assume or hope.