7 ms·
Watch R1 "think" with animated chains of thought
- londons_explore 2y agoThat data looks pretty close to rand() to me...
- ipsum2 2y agoIt seems kinda silly to use a separate service to generate embeddings for t-SNE when you have the embeddings in the model already.
- higuidebot 2y agoIs it generating embeddings or just coordinates? What would be a better way?
- gavmor 2y agoWhat are embeddings if not "just coordinates"?
- higuidebot 2y agoWell ... we have to reduce them to a 2D plane to visualize them ...
- mikeshi42 2y agoSomething needs to generate the document embeddings since the LLM itself won't
- ipsum2 2y agoNo, this is completely wrong. You can get embeddings from the LLM itself, e.g the last layer.
- mikeshi42 2y agoDoesn't the last layer output a variable-size vector based on seq length? It'd take a bit of hacking to get it to be a semantic vector. Additionally, that vector is trained to predict next token as opposed to semantic similarity. I'd assume models trained specifically towards semantic similarity would outperform (I have not bothered comparing both in the past - MMTEB seems to imply so) At that point - it seems quite reasonable to just pass the sentence into an embedding model.
- qoez 2y agoFun experiment but in the back of my mind I suspect this is just plotting a random walk.
- vekntksijdhric 2y agosame, it is unclear what you can get out of this work
- higuidebot 2y agoRandom walk is definitely possible. Also possible that we're observing some "search" in the embedding space from an initial point. It's hard to tell because the chains are often similar lengths, so I don't think it really terminates early. It might be interesting to find the closest CoT component to the final answer and see how step distance inflects at that point
- DiscourseFan 2y agoPerhaps assign each co-ordinate to musical notation and you can get some Scheonberg-esque compositions
- gavmor 2y agoI was, personally, hoping to see a sort of "spiraling down" towards the answer or pathfinding like IDA*[0], but I suppose what we're looking at isn't too dissimilar from A* or Djikstra's if you squint. I suspect you recognize dimensionality reduction, but to reiterate for my own understanding: t-Distributed Stochastic Neighbor Embedding (t-SNE) is one method among a few other, (more popular?) ones like Principal Component Analysis (PCA) and Uniform Manifold Approximation and Projection (UMAP). Is t-SNE the most appropriate technique for modeling the terrain under a multidimensional "walk"? Possibly a linear technique (PCA, LDA, SVD?) or PaCMAP[1], which "dynamically employs a particular set of mid-near pairs to capture the global structure and then improve the local structure." (Qattous H, 2023) 0. https://qiao.github.io/PathFinding.js/visual/ https://qiao.github.io/PathFinding.js/visual/ 1. https://pmc.ncbi.nlm.nih.gov/articles/PMC10756978/ https://pmc.ncbi.nlm.nih.gov/articles/PMC10756978/ Edit: for reference, the tensor projector: https://projector.tensorflow.org/ https://projector.tensorflow.org/
- jejeyyy77 2y agolol isnt this just plotting noise
- higuidebot 2y agoWhether or not it's "noise" might depend on if you think Chains of Thought are causally relevant? I liked these pieces if you want to read more about CoT / O1: O1 Technical Primer: https://www.lesswrong.com/posts/byNYzsfFmb2TpYFPW/o1-a-technical-primer https://www.lesswrong.com/posts/byNYzsfFmb2TpYFPW/o1-a-techn... Using Search Was a Psyop: https://www.interconnects.ai/p/openais-o1-using-search-was-a-psyop https://www.interconnects.ai/p/openais-o1-using-search-was-a... Value Attribution: https://www.lesswrong.com/posts/FX5JmftqL2j6K8dn4/shapley-value-attribution-in-chain-of-thought https://www.lesswrong.com/posts/FX5JmftqL2j6K8dn4/shapley-va...
- atorodius 2y agoIs this using t SNE? Or sth else? I have a feeling similarity is not well defined in whatever space this is using. t SNE is famously unsuited to plot how "close" two points are. it is for clustering
- higuidebot 2y ago2D plot is tSNE, consecutive distance comparison is cosine sim distance normalized across the chain of thought
- vukadinovic 2y agoDistances in t-SNE/UMAP don't mean anything. They are clustering algorithms
- higuidebot 2y ago2D plot is tSNE, consecutive distance comparison is cosine sim distance normalized across the chain of thought
- ThouYS 2y agodoes this show anything?
- jurgenaut23 2y agoAs useless as it gets, surprised that it got to the front page.
- luyu_wu 2y agoIt's a cute little project, arguably far more interesting than the political flame wars that make front page.
- higuidebot 2y agoThank you for your support Mr. Wu
- higuidebot 2y agoI too am pleasantly surprised
- ganyu 2y agoBear in mind that "any two high-dimensional vectors are almost always orthogonal".
- TaurenHunter 2y agoSo the trick is to pick the dimensions that are relevant and discard the rest when calculating the distance.
- frizkie 2y agoIs this better rephrased as “any two vectors in a high-dimensional space are almost always functionally orthogonal”? I have mostly a laypersons understanding of this idea but I would assume that it would be false to say that they are typically _entirely_ orthogonal?
- viraptor 2y agohttps://softwaredoug.com/blog/2022/12/26/surpries-at-hi-dimensions-orthoginality https://softwaredoug.com/blog/2022/12/26/surpries-at-hi-dime... it's both much more likely to be actually orthogonal and almost always very close to orthogonal.
- GeneralMayhem 2y agoThat link doesn't contradict the person you're replying to. Actual orthogonality still has a probability of zero, just as the equator of a sphere has zero surface area, because it's a one-dimensional line (even if it is in some sense "bigger" than the Arctic circle). If you're picking a random point on the (idealized) Earth, the probability of it being exactly on the equator is zero, unless you're willing to add some tolerance for "close enough" in order to give the line some width. Whether that tolerance is +/- one degree of arc, or one mile, or one inch, or one angstrom, you're technically including vectors that aren't perfectly orthogonal to the pole as "successes". That idea does generalize into higher dimensions; the only part that doesn't is the shape of the rest of the sphere (the spinning-top image is actually quite handy).
- 2y ago
- levocardia 2y agoIt would be much more interesting to see PCA (or t-SNE or whatever) on the internal representation within the model itself. As in the activations of a certain number of layers or neurons, as they change from token to token. I don't think the OpenAI embeddings are necessarily an appropriate "map" of the model's internal thoughts. I suppose that raises another questions: Do LLMs "think" in language? Or do they think in a more abstract space, then translate it to language later? My money is on the latter.
- higuidebot 2y agoText embeddings are underused WRT model understanding IMO. "Interpretability" focuses on more complex tools but perhaps misses some of the basics - shouldn't we have some sort of visual understanding of model thinking?
- eightysixfour 2y ago> I suppose that raises another questions: Do LLMs "think" in language? Or do they think in a more abstract space, then translate it to language later? My money is on the latter. The processing happens in latent space and then is converted to tokens/token space. There is research into reasoning models which can spend extra compute in latent space instead of in token space: https://arxiv.org/abs/2412.06769 https://arxiv.org/abs/2412.06769
- HarHarVeryFunny 2y agoI'd have to guess that the "transformations" being made to the embeddings at each layer are basically/mostly just adding (tagging with) incremental levels of additional grammatical/semantic information that has been gleaned by the hierarchical pattern matching that is taking place. At the end of the day our own "thinking" has to be a purely mechanical process, and one also based around pattern recognition and prediction, but "thinking" seems a bit of a loaded term to apply to LLMs given the differences in "cognitive architecture", and smacks a bit of anthromorphism. Reasoning (search-like chained predictions) is more of an algorithmic process, but it seems that the "reactive" pass-thru predictions of the base LLM are more clearly viewed just as pattern recognition and extrapolation/prediction. Prove me wrong!
- antirez 2y agoThe relation among the internal model representations inside its latent space and the embedding of the CoT compressed with a text embedding model is, more or less, minimal. Then we take this and map it to a 2D space, which captures more or less nothing of the original dimentionality and meaning. That's basically plotting random points.
- higuidebot 2y agoPotentially it's useful to understand a model "on its own terms" via its observable outputs. >The relation among the internal model representations inside its latent space and the embedding of the CoT compressed with a text embedding model is, more or less, minimal. This may or may not be correct but one way to find out is by taking a look!
- ekianjo 2y agoKind of useless in 2D space. Not sure what they think they are showing here in the visualization.
- stared 2y agoWhile I like the idea of measuring subsequent steps, this kind of approach of using embeddings is the reason why I wrote: "Don't use cosine distance carelessly" (https://p.migdal.pl/blog/2025/01/dont-use-cosine-similarity https://p.migdal.pl/blog/2025/01/dont-use-cosine-similarity). In this case, cosine distance one would be in a case when it repeats word-by-word. It is not even a "similar thought" but some sort of LLM's OCD. For anything else... cosine similarity says little. Sometimes, two steps can have opposite conclusions but have very high cosine similarity. In another case, it can just expand on the same solution but use different vocabulary or look from another angle. A more robust approach would be to give the whole reasoning to an LLM and ask to grade according to a given criterion (e.g. "grade insight in each step, from 1 to 5").
- higuidebot 2y agoWell, you are certainly correct about how cosine sim would apply to the text embeddings, but I disagree about how useful that application is to our understanding of the model. > In this case, cosine distance one would be in a case when it repeats word-by-word. It is not even a "similar thought" but some sort of LLM's OCD. Observing that would be helpful in our understanding of the model! > For anything else... cosine similarity says little. Sometimes, two steps can have opposite consultation, but they have very high cosine similarity. In another case, it can just expand on the same solution but use different vocabulary or look from another angle. Yes, that would be good to observe also! But here I think you undervalue the specificity of the OAI embeddings model, which has 3072 dimensions. That's quite a lot of information being captured. > A more robust approach would be to give the whole reasoning to an LLM and ask to grade according to a given criterion (e.g. "grade insight in each step, from 1 to 5"). Totally disagree here, using embeddings is much more reliable / robust, I wouldn't put much stock in LLM output, too much going on
- DHolzer 2y ago> Totally disagree here, using embeddings is much more reliable / robust, I wouldn't put much stock in LLM output, too much going on I think both ways can be the preferable option, depending on how well the embedding space represents the text - and that is mostly dependet on the specific use case and model combination. So if the embedding space does not correctly project required nuance, then it's often a viable option to get the top_n results and do the rest by utilizing the llm + validation calls. But i do agree with you, i would always like to work with embeddings rather than some llm output. I think it would be such a great thing to have rock solid embedding space where one would not even consider to look at token predictor models.
- vale95ntino 2y agoPretty cool. Funnily enough I made something similar this weekend that converts CoT to Graphs/Trees of Thoughts with an LLM. It kind of allows to see when the LRM/LLM changes direction or when it finds a path that it wants to follow. Link: https://github.com/vale95ntino/cot2tot https://github.com/vale95ntino/cot2tot
- higuidebot 2y agoVery cool!
- jumploops 2y agoI have the same concern as other commenters (using a separate embedding model, utility of cosine similarity, etc.) BUT this could “seed” a really neat loading graphic for reasoning models, beyond seeing the thinking steps.
- higuidebot 2y agoHa, if steps have consistent distances you could take the average distance at step X and generate a step of that length in some direction and be ~approximately correct regardless of the actual value
- dan_voronov 2y agoWhy not in a three-dimensional space?
- KTibow 2y agoI wonder if you could turn it into a graph by adding connections between the two most similar entries until everything is connected.
- KTibow 2y agoI tried this: https://github.com/KTibow/thoughtgraph https://github.com/KTibow/thoughtgraph
- higuidebot 2y agoLol @ R1's description of my Github ... thanks Deepseek, very cool!
- curiousgal 2y agoLove me some graphs with unlabelled axes...
- higuidebot 2y agoYou're right curiousgal. I filed an issue in response to this comment and will resolve sometime soon: https://github.com/dhealy05/frames_of_mind/issues/1 https://github.com/dhealy05/frames_of_mind/issues/1
- nomilk 2y agoInterestingly, the dots moving about the 2D space resemble eye movements humans involuntarily make when thinking hard and creatively. E.g. looking at random spots (often slightly upwards) and moving to a different place every few seconds or so. Like a kind of synkinesis between brain and eye.
- bowsamic 2y agoFor those of you who were confused about what "R1" is like I was, it seems to be an LLM of some kind https://api-docs.deepseek.com/news/news250120 https://api-docs.deepseek.com/news/news250120
- Sharlin 2y agoNot just "an" LLM, but the LLM (or specifically its chain-of-thought variant) that was in the news for days and caused NVIDIA stocks to crash (well, I'd prefer calling it a natural corrective action by the market) and almost a panic in the US because Chinese and so on.
- rtkwe 2y agoIMO the panic came from a few places: 1) If it's actually much cheaper to build a new model the position of OpenAI, et al is weaker because the barrier against new competitors is much lower than expected. 2) If you don't need a billion dollars in GPUs to build a model NVIDIA maybe isn't worth as much either because you don't need as much compute power to get to the same output. (IMO their valuation is still insanely high but that's a different story) I don't think the drop was entirely irrational they showed there were places a lot of efficiency could be gained in the training process that people didn't seem to be working on in the existing players.
- bowsamic 2y agoFair enough. I'm not interested in this stuff
- brap 2y agoWhenever I hear "chain of thought" I feel like thoughts aren't really linear, it's probably more like "graph of thought". I wonder how models might be able to benefit from structuring their "thinking" this way in a multi-threaded environment
- fudged71 2y agoHmm... a slightly different approach I've seen in the past is that instead of embedding each step separately, you concatenate them together. I wonder if that would make the path more linear/understandable?