5 ms·
I get why everyone is wondering whether this is autogenerated jargon, but I think the basic thrust of this is that machine learning algorithms work better as th
by maxander 6y ago
I get why everyone is wondering whether this is autogenerated jargon, but I think the basic thrust of this is that machine learning algorithms work better as they use larger and larger vectors for representation. I'm pretty sure is just a well-known thing in ML, although the projects listed here may be pushing vector size farther than most.
So for instance, a typical NLP algorithm (although not GPT-3, IIRC) might represent a word as a 500-float-long vector, which is the same as saying the algorithm considers each word as a point in 500-dimensional space. This turns out to have weirdly useful properties, to the point where directions in this 500-dimensional space start to have semantic correspondences (e.g. [0], still one of the coolest things in ML, IMHO.) You can't do the same trick with a 3D space- the algorithm doesn't have enough to work with when all it knows about a word is three numbers.
Another cool example- in gradient descent, you're constantly trying to find the lowest point in a "fitness landscape"; in a 3D landscape, you might easily find yourself in a "valley" where every direction is worse than you currently are (a local minima), and you won't know where to go. In a 500D landscape, it's unlikely that you'll find yourself in a valley where all 500 available directions lead somewhere worse. So the algorithm will be much less likely to get stuck, and this effect gets more robust the more dimensions you have.
[0] https://colah.github.io/posts/2014-07-NLP-RNNs-Representations/ https://colah.github.io/posts/2014-07-NLP-RNNs-Representatio...
- feoren 6y agoThat's a very good explanation of why high-dimensional vectors are so useful, but it doesn't seem to have much to do with this post. How many bits is a 500D vector? Is it 10,000 bits, like in the post? Which "coding" does it use, real? Is it a memory-centric 10,000-bit hyperdimensional real-coded vector that can be combined with permutation? What benefit does your 500D vector get from the fact that it's holographic, not micro-coded? How do you do gradient descent without backpropagation? None of those words make any sense and that's exactly why this post is being lambasted. Thank you for trying to make sense of it, but just because it's possible to write sensibly about similar topics does not mean this post belongs anywhere but the trash can.