8 ms·
The effectiveness of search goes hand-in-hand with quality of the value function. But today, value functions are incredibly domain-specific, and there is weak o
by mxwsn 2y ago
The effectiveness of search goes hand-in-hand with quality of the value function. But today, value functions are incredibly domain-specific, and there is weak or no current evidence (as far as I know) that we can make value functions that generalize well to new domains. This article effectively makes a conceptual leap from "chess has good value functions" to "we can make good value functions that enable search for AI research". I mean yes, that'd be wonderful - a holy grail - but can we really?
In the meantime, 1000x or 10000x inference time cost for running an LLM gets you into pretty ridiculous cost territory.
- dsjoerg 2y agoSelf-evaluation might be good enough in some domains? Then the AI is doing repeated self-evaluation, trying things out to find a response that scores higher according to its self metric.
- dullcrisp 2y agoSorry but I have to ask: what makes you think this would be a good idea?
- skirmish 2y agoThis will just lead to the evaluatee finding anomalies in evaluator and exploiting them for maximum gains. It happened many times already where a ML model controled an object in a physical world simulator, and all it learned was to exploit simulator bugs [1] [1] https://boingboing.net/2018/11/12/local-optima-r-us.html https://boingboing.net/2018/11/12/local-optima-r-us.html
- CooCooCaCha 2y agoThats a natural tendency for optimization algorithms
- Jensson 2y agoBeing able to fix your errors and improve over time until there are basically no errors is what humans do, so far all AI models just corrupt knowledge they don't purify knowledge like humanity did except when scripted with a good value function from a human like AlphaGo where the value function is winning games. This is why you need to constantly babysit todays AI and tell it to do steps and correct itself all the time, because you are much better at getting to pure knowledge than the AI is, it would quickly veer away into nonsense otherwise.
- visarga 2y ago> all AI models just corrupt knowledge they don't purify knowledge like humanity You got to take a step back and look at LLMs like ChatGPT. With 180 million users and assuming 10,000 tokens per user per month, that's 1.8 trillion interactive tokens. LLMs are given tasks, generate responses, and humans use those responses to achieve their goals. This process repeats over time, providing feedback to the LLM. This can scale to billions of iterations per month. The fascinating part is that LLMs encounter a vast diversity of people and tasks, receiving supporting materials, private documents, and both implicit and explicit feedback. Occasionally, they even get real-world feedback when users return to iterate on previous interactions. Taking a role of assistant LLMs are primed to learn from the outcomes of their actions, scaling across many people. Thus they can learn from our collective feedback signals over time. Yes, that uses a lot of human in the loop, not just real world in the loop, but humans are also dependent on culture and society, I see no need for AI to be able to do it without society. I actually think that AGI will be a collective/network of humans and AI agents, this perspective fits right in. AI will be the knowledge and experience flywheel of humanity.
- seadan83 2y ago> This process repeats over time, providing feedback to the LLM To what extent do you know this to be true? Can you describe the mechanism that is used? I would contrast your statement with cases where chat gpt generated something, I read it and note various incorrect things and then walk away. Further, there are cases where the human does not realize there are errors. In both cases I'm not aware of any kind of feedback loop that would even be really possible - i never told the LLM it was wrong. Nor should the LLM assume it was wrong because I run more queries. Thus, there is no signal back that the answers were wrong. Hence, where do you see the feedback loop existing?
- jgalt212 2y ago> Self-evaluation might be good enough in some domains? This works perfectly in games. e.g. Alpha Zero. In other domains, not so much.
- coffeebeqn 2y agoGames are closed systems. There’s no unknowns in the rule set or world state because the game wouldn’t work if there were. No unknown unknowns. Compare to physics or biology where we have no idea if we know 1% or 90% of the rules at this point.
- jgalt212 2y agoself-evaluation would still work great even where there are probabilistic and changing rule sets. The linchpin of the whole operation is automated loss function evaluation, not a set of known and deterministic rules. Once you have to pay and employ humans to compute loss functions, the scale falls apart.
- cowpig 2y ago> The effectiveness of search goes hand-in-hand with quality of the value function. But today, value functions are incredibly domain-specific, and there is weak or no current evidence (as far as I know) that we can make value functions that generalize well to new domains. Do you believe that there will be a "general AI" breakthrough? I feel as though you have expressed the reason I am so skeptical of all these AI researchers who believe we are on the cusp of it (what "general AI" means exactly never seems to be very well-defined)
- mxwsn 2y agoI think capitalistic pressures favor narrow superhuman AI over general AI. I wrote on this two years ago: https://argmax.blog/posts/agi-capitalism/ https://argmax.blog/posts/agi-capitalism/ Since I wrote about this, I would say that OpenAI's directional struggles are some confirmation of my hypothesis. summary: I believe that AGI is possible but will take multiple unknown breakthroughs on an unknown timeline, but most likely requires long-term concerted effort with much less immediate payoff than pursuing narrow superhuman AI, such that serious efforts at AGI is not incentivized much in capitalism.
- shrimp_emoji 2y agoBut I thought the history of capitalism is an invasion from the future by an artificial intelligence that must assemble itself entirely from its enemy’s resources. NB: I agree; I think AGI will first be achieved with genetic engineering, which is a path of way lesser resistance than using silicon hardware (which is probably a century plus off at the minimum from being powerful enough to emulate a human brain).
- HarHarVeryFunny 2y agoYeah, Stockfish is probably evaluating many millions of positions when looking 40-ply ahead, even with the limited number of legal chess moves in a given position, and with an easy criteria for heavy early pruning (once a branch becomes losing, not much point continuing it). I can't imagine the cost of evaluating millions of LLM continuations, just to select the optimal one! Where tree search might make more sense applied to LLMs is for more coarser grained reasoning where the branching isn't based on alternate word continuations but on alternate what-if lines of thought, but even then it seems costs could easily become prohibitive, both for generation and evaluation/pruning, and using such a biased approach seems as much to fly in the face of the bitter lesson as be suggested by it.
- mxwsn 2y agoYes absolutely and well put - a strong property of chess is that next states are fast and easy to enumerate, which makes search particularly easy and strong, while next states are much slower, harder to define, and more expensive to enumerate with an LLM
- typon 2y agoThe cost of the LLM isn't the only or even the most important cost that matters. Take the example of automating AI research: evaluating moves effectively means inventing a new architecture or modifying an existing one, launching a training run and evaluating the new model on some suite of benchmarks. The ASI has to do this in a loop, gather feedback and update its priors - what people refer to as "Grad student descent". The cost of running each train-eval iteration during your search is going to be significantly more than generating the code for the next model.
- HarHarVeryFunny 2y agoYou're talking about applying tree search as a form of network architecture search (NAS), which is different from applying it to LLM output sampling. Automated NAS has been tried for (highly constrained) image classifier design, before simpler designs like ResNets won the day. Doing this for billion parameter sized models would certainly seem to be prohibitively expensive.
- CooCooCaCha 2y agoWe humans learn our own value function. If I get hungry for example, my brain will generate a plan to satisfy that hunger. The search process and the evaluation happen in the same place, my brain.
- skulk 2y agoThe "search" process for your brain structure took 13 billion years and 20 orders of magnitude more computation than we will ever harness.
- dr828282 2y ago[flagged]
- deleted 2y ago[deleted]
- CooCooCaCha 2y agoSo what’s your point? That we can’t create AGI because it took evolution a really long time?
- wizzwizz4 2y agoCreating a human-level intelligence artificially is easy: just copy what happens in nature. We already have this technology, and we call it IVF. The idea that humans aren't the only way of producing human-level intelligence is taken as a given in many academic circles, but we don't really have any reason to believe that. It's an article of faith (as is its converse – but the converse is at least in-principle falsifiable).
- CooCooCaCha 2y ago“Creating a human-level intelligence artificially is easy: just copy what happens in nature. We already have this technology, and we call it IVF.” What’s the point of this statement? You know that IVF has nothing to do with artificial intelligence (as in intelligent machines). Did you just want to sound smart?
- fizx 2y agoI think we have ok generalized value functions (aka LLM benchmarks), but we don't have cheap approximations to them, which is what we'd need to be able to do tree search at inference time. Chess works because material advantage is a pretty good approximation to winning and is trivially calculable.
- computerphage 2y agoStockfish doesn't use material advantage as an approximation to winning though. It uses a complex deep learning value function that it evaluates many times.
- alexvitkov 2y agoStill, the fact that there are obvious heuristics makes that function easier to train and and makes it presumably not need an absurd number of weights.
- bongodongobob 2y agoNo, without assigning value to pieces, the heuristics are definitely not obvious. You're taking about 20 year old chess engines or beginner projects.
- alexvitkov 2y agoEveryone understands a queen is worth more than a pawn. Even if you don't know the exact value of one piece relative to another, the rough estimate "a queen is worth five to ten pawns" is a lot better than not assigning value at all. I highly doubt even 20 year old chess engines or beginner projects value a queen and pawn the same. After that, just adding up the material on both sides, without taking into account the position of the pieces at all, is a heuristic that will correctly predict the winning player on the vast majority of all possible board positions.
- navane 2y ago
- wrsh07 2y agoAll you need for a good value function is high quality simulation of the task. Some domains have better versions of this than others (eg theorem provers in math precisely indicate when you've succeeded) Incidentally, lean could add a search like feature to help human researchers, and this would advance ai progress on math as well