8 ms·
I feel it is interesting but not what would be ideal. I really think if the models could be less linear and process over time in latent space you'd get somethin
by robviren 1y ago
I feel it is interesting but not what would be ideal. I really think if the models could be less linear and process over time in latent space you'd get something much more akin to thought. I've messed around with attaching reservoirs at each layer using hooks with interesting results (mainly over fitting), but it feels like such a limitation to have all model context/memory stuck as tokens when latent space is where the richer interaction lives. Would love to see more done where thought over time mattered and the model could almost mull over the question a bit before being obligated to crank out tokens. Not an easy problem, but interesting.
- dkersten 1y agoAgree! I’m not an AI engineer or researcher, but it always struck me as odd that we would serialise the 100B or whatever parameters of latent space down to maximum 1M tokens and back for every step.
- vonneumannstan 1y ago>I feel it is interesting but not what would be ideal. I really think if the models could be less linear and process over time in latent space you'd get something much more akin to thought. Please stop, this is how you get AI takeovers.
- varelse 1y ago[dead]
- adastra22 1y agoCitation seriously needed.
- vonneumannstan 1y agoIt's really very simple. As models become more capable they may become interested in deceiving humans or otherwise manipulating them to achieve their goals. We already see this in various places see: https://www.anthropic.com/research/agentic-misalignment https://www.anthropic.com/research/agentic-misalignment https://arxiv.org/abs/2412.14093 https://arxiv.org/abs/2412.14093 If the chain of thought of models becomes pure "neuralese" i.e. the models think purely in latent space then we will lose the ability to monitor for malicious behavior. This is incredibly dangerous, CoT monitoring is one of the best and highest leverage tools for monitoring model behavior and losing that would be devastating for safety. https://www.lesswrong.com/posts/D2Aa25eaEhdBNeEEy/worries-about-latent-reasoning-in-llms https://www.lesswrong.com/posts/D2Aa25eaEhdBNeEEy/worries-ab... https://www.lesswrong.com/posts/mpmsK8KKysgSKDm2T/the-most-forbidden-technique https://www.lesswrong.com/posts/mpmsK8KKysgSKDm2T/the-most-f... https://www.lesswrong.com/posts/3W8HZe8mcyoo4qGkB/an-idea-for-avoiding-neuralese-architectures-1 https://www.lesswrong.com/posts/3W8HZe8mcyoo4qGkB/an-idea-fo... https://x.com/RyanPGreenblatt/status/1908298069340545296 https://x.com/RyanPGreenblatt/status/1908298069340545296 https://redwoodresearch.substack.com/p/notes-on-countermeasures-for-exploration https://redwoodresearch.substack.com/p/notes-on-countermeasu...
- CuriouslyC 1y agoThey're already implementing branching thought and taking the best one, eventually the entire response will be branched, with branches being spawned and culled by some metric over the lifetime of the completion. It's just not feasible now for performance reasons.