4 ms·
Harnessing the Universal Geometry of Embeddings
- measurablefunc 27d agoWhat is the (co)homology of this space?
- chombier 26d agoThat of the underlying, hypothetical universal brain topology?
- measurablefunc 26d agoAnd what would that be?
- nickledave 27d agoDupe: https://news.ycombinator.com/item?id=44054425 https://news.ycombinator.com/item?id=44054425 Note this is version 4 of the paper and the original post was version 1 (I think?) OpenReview (for NeurIPS) for the curious: https://openreview.net/forum?id=jiCLUPq5xv https://openreview.net/forum?id=jiCLUPq5xv
- srean 27d agoLet's assume that monotonocity of pair-wise distances are preserved. Without knowing the details of how the paper solved the problem, my first attempt would be to find the diametrically distant pair of points in the two different embeddings and assume that the pair is the same pair. Then find the next distant pairs and so on. After sufficiently many such pairs have been found, or better still, the largest d-simplex is found, find that scaled rigid body transformation that makes the corresponding pairs coincide. Proceeding this way ought to be less work than solving a generic graph isomorphism problem.
- robrenaud 26d agoI think a less stringent, but still workable assumption is that for very similair objects, their distances will be small. This is much easier to accomplish than agreement across all pairs.
- srean 26d agoCould you explain a bit more. What you say about similar objects is obviously true. However the algorithm sketch that you have in your mind is a little implicit. Could you make it more explicit. I am quite curious. I explained my thoughts in a comment here https://news.ycombinator.com/item?id=49595424 https://news.ycombinator.com/item?id=49595424
- ironSkillet 27d agoI am not familiar with the standards of publishing in machine learning, but as someone trained in a mathematics background, this paper seems relatively light on details and heavy on exposition. Is that typical? Is this a really novel idea? Not trying to be snarky, just trying to understand how meaningful this is.
- rhelz 27d agoYou are not wrong. But this has by no means proven its up to the standard of being publishable in a machine learning journal. Its on arXiv.org, which, lets face it, at the end of the day is a vanity press.
- efavdb 27d agoAt a minimum posting to arxiv gives others a standard way to cite the work.
- odyssey7 27d agoThe pace of things is moving along so rapidly right now, I’m not sure that waiting for peer reviews is always a wise move. Doubly so if there’s a paywall; why limit your article’s impact by placing it where practitioners’ agents might not be able to access it? The rapid progress right now is challenging for conventional academic processes. If the value of the paper is difficult to independently verify, for example, if it depends on the credibility of the author, then the academic ritual can add something. If it’s a mathematical result, one that can be automatically verified, or a machine learning technique that anyone can try with Claude code reconstructing it for them, this sort of pre-print publishing model is advantageous.
- rhelz 27d ago// why limit your article's impact // Because...science? It's not science until it passes peer review. I'm not advocating that everybody stops posting to arXiv, and I'm not saying you can't find good stuff there. I'm just saying, it's a vanity press, there is absolutely no guarantee of the paper's quality. And being published by a famous professor from a prestigious university is also no guarantee. If we've learned anything from the non-reproducibility crisis, it is that a paper's origin story is no guarantee.
- rhelz 27d agoCyberphrenology. In any two random graphs, you'll find an isomorphic graph which is can be up to log of the size of the graphs. And if the LLM has been trained up to the limit of what data it can hold, it is going to be random. Proof below if it isn't obvious. The entire effort of all people who are trying to understand how LLMs work, how they represent their data, its all bound to fail. Proof: a LLM is a very good approximation of the Solomonov/Levin/Kolmogorov universal probability function on tokens. As such, it will be random--pure white noise--because if you found any patterns in there, you could exploit the regularity and come up with a smaller set of weights for the same LLM. There are no patterns there to be found. They have all been factored out by training the neural net until it couldn't learn any more.
- sdenton4 27d ago/a smaller set of weights for the same LLM./ Distillation is alive and well... Earlier work on model printing also found that it's pretty easy to find smaller sets of parameters which can replicate the behavior of the entire network with pretty good fidelity. Large parameter counts give space to explore, and give routes out of what would be local minima in a lower dimensional space. In other words, there's no guarantee that any given trained model is a minimal representation of its training set.
- rhelz 27d agoI'm not claiming any arbitrary set of weights is a minimal representation. But typically, if people could achieve the same quality of results with a smaller set of weights, or weights which have been quantized to lower bit representations, etc, they would have published the smaller one instead.
- andrewflnr 26d agoYou kind of are claiming they're minimal, though. Because if they're not, your statement that "if you found any patterns in there, you could exploit the regularity..." implies nothing. Yeah, the patterns are there, and people are exploiting them. Your socioeconomic argument just doesn't hold either. People don't delay releasing models until they've minimized it to the theoretical limit. They ship it when it's good enough for whatever job they're making it for.
- paidx 27d ago[flagged]
- stephantul 26d agoI’ve never liked that this was called “the platonic representation hypothesis”. Lots of weird baggage attached and seems like a waste of a good name.
- pksebben 26d ago"We believe these representations are not serious, they're just really good friends."
- srean 26d agoOne way to pose/(think about) the problem is that there are two finite metric spaces linked by an unknown odometry (damn you autocorrect). The problem is to recover that unknown isometry. This, like graph isometry, can be very computationally intensive in the worst case. However, heuristics to aid matching one vertex on one graph to another vertex on another graph using local, semilocal structural signatures can be very effective on particular cases. One can of course argue that the spaces are not designed as metric spaces. Even if true, these might be metrizable topological spaces. More generally, if these are indeed non-metric spaces one can still pose it as finding the unknown isomorphism between two poset spaces. In my other comment I was using the property of maximal chains -- Identify the longest chains in both posets. The isomorphism must map the longest chain in Poset 1 directly to a longest chain in Poset 2, preserving the exact linear order.
- fennecfoxy 26d agoI'm not as heavy on the maths stuff involved in this as other people commenting appear to be. But the idea makes sense, of course there is still recoverable data in embeddings, that's the point. Though as I constantly find the more you try to squeeze into an n bit vector the more watered down everything gets. I suppose a latent space could be encrypted/mapped in some way to resolve that, but how many people are exposing their vectors in the first place?
- ViscountPenguin 26d agoThe point of the paper isn't that embeddings contain information, it's that even if you don't know what model generated a set of embeddings you can still recover information from the geometry of the point cloud itself. The fact that this is possible also adds some pretty strong restriction to the set of possible maps you could use to remove that information. No linear map will work since all embedding spaces are ~an orthonormal matrix apart, so some form of encryption is necessary. This wasn't known until very recently.
- fennecfoxy 25d agoAh right, thank you for clarifying. But I think my point still stands, isn't the geometry information THE information I referred to in the first place? Obviously the vector size gives you the granularity but it's kind of unavoidable to positionally encode information in a latent space...that's literally what they're for? But yes, it is very cool to know that regardless of exact implementation finding x,y,z representations of some dataset with various relationships (like language) creates similar geometry/clues across all the implementations.
- dankai 26d agoI've been working on research related to this in the context of diferent LLMs, and I can tell you that similarity != executability. While you can make embeddings across different LLMs similar (i.e. universally looking) the few percentage of R2 that you are missing in translating the universal representation into a native representation are precisely those that make the hidden states executable in the LLM (which makes them useful). They do this with simple embedding models because there the purpose of the embeddings is to measure similarity, but if you would like to use this principle to turn latent representations of one LLM into latent representations that are understandable/executably by a different LLM, you will fail.
- one-bank4326 26d ago[flagged]
- FloorEgg 26d agoThis is being positioned as evidence of the Platonic Representation Hypothesis, but isn't it more likely an artifact of training models on the same/similar data sets?
- ahmedelsama 26d ago[flagged]