3 ms·
Far too much marketing speech, far too little math or theory, and completely misses the mark on the 'next frontier'. Maybe four years ago, spatial reasoning was
by programjames 11mo ago
Far too much marketing speech, far too little math or theory, and completely misses the mark on the 'next frontier'. Maybe four years ago, spatial reasoning was the problem to solve, but by 2022 it was solved. All that remained was scaling up. The actual three next problems to solve (in order of when they will be solved) are:
- Reinforcement Learning (2026)
- General Intelligence (2027)
- Continual Learning (2028)
EDIT: lol, funny how the idiots downvote
- 7moritz7 11mo agoHasn't RLHF and with LLM feedback been around for years now
- programjames 11mo agoLarge latent flow models are unbiased. On the other hand, if you purely use policy optimization, RLHF will be biased towards short horizons. If you add in a value network, the value has some bias (e.g. MSE loss on the value --> Gaussian bias). Also, most RL has some adversarial loss (how do you train your preference network?), which makes the loss landscape fractal which SGD smooths incorrectly. So, basically, there's a lot of biases that show up in RL training which can make it both hard to train, and even if successful, not necessarily optimizing what you want.
- storus 11mo agoWe might not even need RL as DPO has shown.
- programjames 11mo ago> if you purely use policy optimization, RLHF will be biased towards short horizons > most RL has some adversarial loss (how do you train your preference network?), which makes the loss landscape fractal which SGD smooths incorrectly
- koakuma-chan 11mo agoIn my thinking what AI lacks is a memory system
- 7moritz7 11mo agoThat has been solved with RAG, OCR-ish image encoding (deepseek recently) and just long context windows in general.
- koakuma-chan 11mo agoNot really. For example we still can’t get coding agents to work reliably, and I think it’s a memory problem, not a capabilities problem.
- atlex2 11mo agoOn the other hand, test-time weight updates would make model interpretability much harder.
- Eisenstein 11mo agoRAG is like constantly reading your notes instead of integrating experiences into your processes.
- deleted 11mo ago[deleted]
- l9o 11mo agoWhat do you consider "General Intelligence" to be?
- programjames 11mo agoA good start would be: 1. Robust to adversarial attacks (e.g. in classification models or LLM steering). 2. Solving ARC-AGI. Current models are optimized to solve the current problem they're presented, not really find the most general problem-solving techniques.
- stirfish 11mo agoI like to think I'm generally intelligent, but I am not robust to adversarial attacks. Edit: I'm trying arc-agi tests now and it's looking bad for me: https://arcprize.org/play?task=e3721c99 https://arcprize.org/play?task=e3721c99
- programjames 11mo ago"I like to think I'm generally intelligent, but I am not robust to adversarial attacks. I'm trying arc-agi tests now and it's looking bad for me." One man's modus ponens is another man's modus tollens. "I'm trying arc-agi tests now and it's looking bad for me. I am not robust to adversarial attacks. I think I'm not generally intelligent."
- stirfish 11mo agoI forgot what modus ponens/tollens are, but you get get it - I think I'm not generally intelligent For people coming after me, or for anyone who took discrete math a decade ago and need a quick refresher: Modus ponens (affirming): if P, then Q. P is true, therefore Q. If it is raining, the grass is wet. It is raining. Therefore the grass is wet. Modus tollens (denying): if P, then Q. Q is false. Therefore P is false. If it is raining, then the grass is wet. The grass is not wet. Therefore, it is not raining.
- whatever1 11mo agoCombinatorial search is also a solved problem. We just need a couple of Universes to scale it up.
- programjames 11mo agoIf there isn't a path humans know how to take with their current technology, it isn't a solved problem. It's much different than people training an image model for research purposes, and knowing that $100m in compute is probably enough for a basic video model.