3 ms·
I've probably read 100 ML papers. They're pretty bad. Even the most cited ones like Adam, Attention, and Latent Diffusion are messy, unclear, or just straight u
by programjames 3y ago
I've probably read 100 ML papers. They're pretty bad. Even the most cited ones like Adam, Attention, and Latent Diffusion are messy, unclear, or just straight up wrong at times.
Adam - Their algorithm is inefficient. Their paper can be summarized as "use the signal-to-noise ratio," but the key idea is hidden in a page-long paragraph halfway through the paper.
Attention - I probably read "keys, query, values" a dozen times before I realized the key and query matrices multiply with the same vector (where the "self" comes in).
Latent Diffusion - Their theoretical grounding was just completely wrong. It's not some stochastic diffeq, or Markov process, or w/e terms they threw around. It's just finding the gradient of log-likelihood using finite differences (the error they add).