14 ms·
AI Search: The Bitter-Er Lesson
- johnthewise 2y agoWhat happened to all the chatter about Q*? I remember reading about this train/test time trade-off back then, does anyone have good list of recent papers/blogs about this? What is holding back this or openai is just running some model 10x longer to estimate what they would get if they trained with 10x? This tweet is relevant: https://x.com/polynoamial/status/1676971503261454340 https://x.com/polynoamial/status/1676971503261454340
- Kronopath 2y agoAnything that allows AI to scale to superinteligence quicker is going to run into AI alignment issues, since we don’t really know a foolproof way of controlling AI. With the AI of today, this isn’t too bad (the worst you get is stuff like AI confidently making up fake facts), but with a superintelligence this could be disastrous. It’s very irresponsible for this article to advocate and provide a pathway to immediate superintelligence (regardless of whether or not it actually works) without even discussing the question of how you figure out what you’re searching for, and how you’ll prevent that superintelligence from being evil.
- nullc 2y agoI don't think your response is appropriate. Narrow domain "superintelligence" is around us everywhere-- every PID controller can drive a process to its target far beyond any human capability. The obvious way to incorporate good search is to have extremely fast models that are being used in the search interior loop. Such models would be inherently less general, and likely trained on the specific problem or at least domain-- just for performance sake. The lesson in this article was that a tiny superspecialized model inside a powerful transitional search framework significantly outperformed a much larger more general model. Use of explicit external search should make the optimization system's behavior and objective more transparent and tractable than just sampling the output of an auto-regressive model alone. If nothing else you can at least look at the branches it did and didn't explore. It's also a design that's more easy to bolt in varrious kinds of regularizes, code to steer it away from parts of the search space you don't want it operating in. The irony of all the AI scaremongering is that if there is ever some evil AI with some LLM as an important part of its reasoning process if it is evil it may well be so because being evil is a big part of the narrative it was trained on. :D
- coldtea 2y agoOf course "superintelligence" is just a mythical creature at the moment, with no known path to get there, or even a specific proof of what it even means - usually it's some hand waving about capabilities that sound magical, when IQ might very well be subject to diminishing returns.
- drdeca 2y agoDo you mean no way to get there within realistic computation bounds? Because if we allow for arbitrarily high (but still finite) amounts of compute, then some computable approximation of AIXI should work fine.
- coldtea 2y ago>Do you mean no way to get there within realistic computation bounds? I mean there's no well defined "there" either. It's a hand-waved notion that adding more intelligence (itself not very well defined, but let's use IQ) you get to something called "hyperintelligence", say IQ 1000 or IQ 10000, that has what can be described as magical powers, like it can convince any person to do anything, can invent things at will, huge business success, market prediction, and so on. Whether intelligence is cummulative like that, or whether having it gets you those powers (aside from the succesful high IQ people, we know many people with IQ 145+ that are not inventing stuff left and right, or convincing people with some greater charisma than the average IQ 100 or 120 politician, but e.g. are just sad MENSA losers, whose greatest achievement is their test scores). >Because if we allow for arbitrarily high (but still finite) amounts of compute, then some computable approximation of AIXI should work fine. I doubt that too. The limit for LLMs for example is more human produced training data (a hard limit) than compute.
- drdeca 2y ago> itself not very well defined, but let's use IQ IQ has an issue that is inessential to the task at hand, which is how it is based on a population distribution. It doesn’t make sense for large values (unless there is a really large population satisfying properties that aren’t satisfied). > I doubt that too. The limit for LLMs for example is more human produced training data (a hard limit) than compute. Are you familiar with what AIXI is? When I said “arbitrarily large”, it wasn’t for laziness reasons that I didn’t give an amount that is plausibly achievable. AIXI is kind of goofy. The full version of AIXI is uncomputable (it uses a halting oracle), which is why I referred to the computable approximations to it. AIXI doesn’t exactly need you to give it a training set, just put it in an environment where you give it a way to select actions, and give it a sensory input signal, and a reward signal. Then, assuming that the environment it is in is computable (which, recall, AIXI itself is not), its long-run behavior will maximize the expected (time discounted) future reward signal. There’s a sense in which it is asymptotically optimal across computable environments (... though some have argued that this sense relies on a distribution over environments based on the enumeration of computable functions, and that this might make this property kinda trivial. Still, I’m fairly confident that it would be quite effective. I think this triviality issue is mostly a difficulty of having the right definition.) (Though, if it was possible to implement practically, you would want to make darn sure that the most effective way for it to make its reward signal high would be for it to do good things and not either bad things or to crack open whatever system is setting the reward signal in order for it to set it itself.) (How it works: AIXI basically enumerates through all possible computable environments, assigning initial probability to each according to the length of the program, and updating the probabilities based on the probability of that environment providing it with the sequence of perceptions and reward signals it has received so far when the agent takes the sequence of actions it has taken so far. It evaluates the expected values of discounted future reward of different combinations of future actions based on its current assigned probability of each of the environments under consideration, and selects its next action to maximize this. I think the maximum length of programs that it considers as possible environments increases over time or something, so that it doesn’t have to consider infinitely many at any particular step.)
- aidan_mclau 2y agoHey! Essay author here. >The cool thing about using modern LLMs as an eval/policy model is that their RLHF propagates throughout the search. >Moreover, if search techniques work on the token level (likely), their thoughts are perfectly interpretable. I suspect a search world is substantially more alignment-friendly than a large model world. Let me know your thoughts!
- Tepix 2y agoYour webpage is broken for me. The page appears briefly, then there's a french error message telling me that an error occured and i can retry. Mobile Safari, phone set to french.
- mxwsn 2y agoThe effectiveness of search goes hand-in-hand with quality of the value function. But today, value functions are incredibly domain-specific, and there is weak or no current evidence (as far as I know) that we can make value functions that generalize well to new domains. This article effectively makes a conceptual leap from "chess has good value functions" to "we can make good value functions that enable search for AI research". I mean yes, that'd be wonderful - a holy grail - but can we really? In the meantime, 1000x or 10000x inference time cost for running an LLM gets you into pretty ridiculous cost territory.
- dsjoerg 2y agoSelf-evaluation might be good enough in some domains? Then the AI is doing repeated self-evaluation, trying things out to find a response that scores higher according to its self metric.
- dullcrisp 2y agoSorry but I have to ask: what makes you think this would be a good idea?
- skirmish 2y agoThis will just lead to the evaluatee finding anomalies in evaluator and exploiting them for maximum gains. It happened many times already where a ML model controled an object in a physical world simulator, and all it learned was to exploit simulator bugs [1] [1] https://boingboing.net/2018/11/12/local-optima-r-us.html https://boingboing.net/2018/11/12/local-optima-r-us.html
- CooCooCaCha 2y agoThats a natural tendency for optimization algorithms
- Jensson 2y agoBeing able to fix your errors and improve over time until there are basically no errors is what humans do, so far all AI models just corrupt knowledge they don't purify knowledge like humanity did except when scripted with a good value function from a human like AlphaGo where the value function is winning games. This is why you need to constantly babysit todays AI and tell it to do steps and correct itself all the time, because you are much better at getting to pure knowledge than the AI is, it would quickly veer away into nonsense otherwise.
- fire_lake 2y agoI didn’t understand this piece. What do they mean by using LLMs with search? Is this simply RAG?
- roca 2y agoThey mean something like the minmax algorithm used in game engines.
- Legend2440 2y ago“Search” here means trying a bunch of possibilities and seeing what works. Like how a sudoku solver or pathfinding algorithm does search, not how a search engine does.
- fire_lake 2y agoBut the domain of “AI Research” is broad and imprecise - not simple and discrete like chess game states. What is the type of each point in the search space for AI Research?
- moffkalast 2y agoWell if we knew how to implement it, then we'd already have it eh?
- fire_lake 2y agoIn chess we know how to describe all possible board states and the transitions (the next moves). We just don’t know which transition is the best to pick, hence it’s a well defined search problem. With AI Research we don’t even know the shape of the states and transitions, or even if that’s an appropriate way to think about things.
- tsaixingwei 2y agoGiven the example of Pfizer in the article, I would tend to agree with you that ‘search’ in this context means augmenting GPT with RAG of domain specific knowledge.
- timfsu 2y agoThis is a fascinating idea - although I wish the definition of search in the LLM context was expanded a bit more. What kind of search capability strapped onto current-gen LLMs would give them superpowers?
- gwd 2y agoI think what may be confusing is that the author is using "search" here in the AI sense, not in the Google sense: that is, having an internal simulator of possible actions and possible reactions, like Stockfish's chess move search (if I do A, it could do B C or D; if it does B, I can do E F or G, etc). So think about the restrictions current LLMs have: * They can't sit and think about an answer; they can "think out loud", but they have to start talking, and they can't go back and say, "No wait, that's wrong, let's start again." * If they're composing something, they can't really go back and revise what they've written * Sometimes they can look up reference material, but they can't actually sit and digest it; they're expected to skim it and then give an answer. How would you perform under those circumstances? If someone were to just come and ask you any question under the sun, and you had to just start talking, without taking any time to think about your answer, and without being able to say "OK wait, let me go back"? I don't know about you, but there's no way I would be able to perform anywhere close to what ChatGPT 4 is able to do. People complain that ChatGPT 4 is a "bullshitter", but given its constraints that's all you or I would be in the same situation -- but it's already way, way better than I could ever be. Given its limitations, ChatGPT is phenomenal. So now imagine what it could do if it were given time to just "sit and think"? To make a plan, to explore the possible solution space the same way that Stockfish does? To take notes and revise and research and come back and think some more, before having to actually answer? Reading this is honestly the first time in a while I've believed that some sort of "AI foom" might be possible.
- cbsmith 2y ago> They can't sit and think about an answer; they can "think out loud", but they have to start talking, and they can't go back and say, "No wait, that's wrong, let's start again." I mean, technically, they could say that.
- 1024core 2y agoThe problem with adding "search" to a model is that the model has already seen everything to be "search"ed in its training data. There is nothing left. Imagine if Leela (author's example) had been trained on every chess board position out there (I know it's mathematically impossible, but bear with me for a second). If Leela had been trained on every board position, it may have whupped Stockfish. So, adding "search" to Leela would have been pointless, since it would have seen every board position out there. Today's LLMs are trained on every word ever written on the 'net, every word ever put down in a book, every word uttered in a video on Youtube or a podcast.
- deleted 2y ago[deleted]
- yousif_123123 2y agoStill, similar to when you have read 10 textbooks, if you are answering a question and have access to the source material, it can help you in your answer.
- groby_b 2y agoYou're omitting the somewhat relevant part of recall ability. I can train a 50 parameter model on the entire internet, and while it's seen it all, it won't be able to recall it. (You can likely do the same thing with a 500B model for similar results, though it's getting somewhat closer to decent recall) The whole point of deep learning is that the model learns to generalize. It's not to have a perfect storage engine with a human language query frontend.
- sebastos 2y agoFully agree, although it’s interesting to consider the perspective that the entire LLM hype cycle is largely built around the question “what if we punted on actual thinking and instead just tried to memorize everything and then provide a human language query frontend? Is that still useful?” Arguably it is (sorta), and that’s what is driving this latest zeitgeist. Compute had quietly scaled in the background while we were banging our heads against real thinking, until one day we looked up and we still didn’t have a thinking machine, but it was now approximately possible to just do the stupid thing and store “all the text on the internet” in a lookup table, where the keys are prompts. That’s… the opposite of thinking, really, but still sometimes useful! Although to be clear I think actual reasoning systems are what we should be trying to create, and this LLM stuff seems like a cul-de-sac on that journey.
- groby_b 2y agoWhile I respect the power of intuition - this may well be a great path - it's worth keeping in mind that this is currently just that. A hunch. Leela got crushed due to AI directed search, what if we can wave a wand and hand all AIs search. Somehow. Magically. Which will then somehow magically trounce current LLMs at domain-specific task. There's a kernel of truth in there. See the papers on better results via monte carlo search trees (e.g. [1]). See mixture-of-LoRA/LoRA-swarm approaches. (I swear there's a startup using the approach of tons of domain-specific LoRAs, but my brain's not yielding the name) Augmenting LLM capabilities via _some_ sort of cheaper and more reliable exploration is likely a valid path. It's not GPT-8 next year, though. [1] https://arxiv.org/pdf/2309.03224 https://arxiv.org/pdf/2309.03224
- hartator 2y agoIsn't the "search" space infinite though and impossible to qualify "success"? You can't just give LLMs infinite compute time and expect them to find answers for like "cure cancer". Even chess moves that seem finite and success quantifiable are an also infinite problem and the best engines take "shortcuts" in their "thinking". It's impossible to do for real world problems.
- cpill 2y agothe recent episode of Machine Learning Street Talk on control theory for LLMs sounds like it's thinking in this direction. Say you have 100k agents searching through research papers, and then trying every combination of them, 100k^2, to see if there is any synergy of ideas, and you keep doing this for all the successful combos... some of these might give the researchers some good ideas to try out. I can see it happening, if they can fine tune a model that becomes good at idea synergy. but then again real creativity is hard
- _lyxd 2y agoHow would one finetune for "idea synergy"?
- cpill 2y agoYou'd need an example set, like all FT'ing
- _lyxd 2y agoIdea synergy is Ill defined, creating a consistent, high quality dataset of the scale required is not easy at all. What is the data, pairs of syngeristic papers? Or are you embedding each paper and finding the closest in "idea space"? LLMs are apt for RAG/fuzzy search. Learning a relevant embedding space that has the fidelity where distances have an interpretation of "idea symergy" is not like "semantic meaning" (I.e. RAG embedding space), otherwise you'd simply get similar papers on the same topics as "synergistic" Maybe I'm confused as to what you mean by idea synergy.
- salamo 2y agoSearch is almost certainly necessary, and I think the trillion dollar cluster maximalists probably need to talk to people who created superhuman chess engines that now can run on smartphones. Because one possibility is that someone figures out how to beat your trillion dollar cluster with a million dollar cluster, or 500k million dollar clusters. On chess specifically, my takeaway is that the branching factor in chess never gets so high that a breadth-first approach is unworkable. The median branching factor (i.e. the number of legal moves) maxes out at around 40 but generally stays near 30. The most moves I have ever found in any position from a real game was 147, but at that point almost every move is checkmate anyways. Creating superhuman go engines was a challenge for a long time because the branching factor is so much larger than chess. Since MCTS is less thorough, it makes sense that a full search could find a weakness and exploit it. To me, the question is whether we can apply breadth-first approaches to larger games and situations, and I think the answer is clearly no. Unlike chess, the branching factor of real-world situations is orders of magnitude larger. But also unlike chess, which is highly chaotic (small decisions matter a lot for future state), most small decisions don't matter. If you're flying from NYC to LA, it matters a lot if you drive or fly or walk. It mostly doesn't matter if you walk out the door starting with your left foot or your right. It mostly doesn't matter if you blink now or in two seconds.
- cpill 2y agoI think the branching factor for LLMs is around 50k for the number of next possible tokens.
- refulgentis 2y ago100%, GPT-3 <= x < GPT-4o, 100,064, x = GPT-4o, 199,996. (My EoW emergency was the const Map that stored them broke the build, so these #s happen to be top of mind)
- kippinitreal 2y agoI wonder if in an application you could branch on something more abstract than tokens. While there might by 50k token branches and 1k of reasonable likelihood, those actually probably cluster into a few themes you could branch off of. For example “he ordered a …” [burger, hot dog, sandwich: food] or [coke, coffee, water: drinks] or [tennis racket, bowling ball, etc: goods].
- optimalsolver 2y agoCharlie Steiner pointed this out 5 years ago on Less Wrong: >If you train GPT-3 on a bunch of medical textbooks and prompt it to tell you a cure for Alzheimer's, it won't tell you a cure, it will tell you what humans have said about curing Alzheimer's ... It would just tell you a plausible story about a situation related to the prompt about curing Alzheimer's, based on its training data. Rather than a logical Oracle, this image-captioning-esque scheme would be an intuitive Oracle, telling you things that make sense based on associations already present within the training set. >What am I driving at here, by pointing out that curing Alzheimer's is hard? It's that the designs above are missing something, and what they're missing is search. I'm not saying that getting a neural net to directly output your cure for Alzheimer's is impossible. But it seems like it requires there to already be a "cure for Alzheimer's" dimension in your learned model. The more realistic way to find the cure for Alzheimer's, if you don't already know it, is going to involve lots of logical steps one after another, slowly moving through a logical space, narrowing down the possibilities more and more, and eventually finding something that fits the bill. In other words, solving a search problem. >So if your AI can tell you how to cure Alzheimer's, I think either it's explicitly doing a search for how to cure Alzheimer's (or worlds that match your verbal prompt the best, or whatever), or it has some internal state that implicitly performs a search. https://www.lesswrong.com/posts/EMZeJ7vpfeF4GrWwm/self-supervised-learning-and-agi-safety?commentId=vRfNdRe8Gzz9QhFYq https://www.lesswrong.com/posts/EMZeJ7vpfeF4GrWwm/self-super...
- lucb1e 2y agoGeneralizing this (doing half a step away from GPT-specifics), would it be true to say the following? "If you train your logic machine on a bunch of medical textbooks and prompt it to tell you a cure for Alzheimer's, it won't tell you a cure, it will tell you what those textbooks have said about curing Alzheimer's." Because I suspect not. GPT seems mostly limited to regurgitating+remixing what it read, but other algorithms with better logic could be able to essentially do a meta study: take the results from all Alzheimer's experiments we've done and narrow down the solution space to beyond what humans achieved so far. A human may not have the headspace to incorporate all relevant results at once whereas a computer might Asking GPT to "think step by step" helps it, so clearly it has some form of this necessary logic, and it also performs well at "here's some data, transform it for me". It has limitations in both how good its logic is and the window across which it can do these transformations (but it can remember vastly more data from training than from the input token window, so perhaps that's a partial workaround). Since it does have both capabilities, it does not seem insurmountable to extend it: I'm not sure we can rule out that an evolution of GPT can find Alzheimer's cure within existing data, let alone a system even more suited to this task (still far short of needing AGI) This requires the data to contain the necessary building blocks for a solution, but the quote seems to dismiss the option altogether even if the data did contain all information (but not yet the worked-out solution) for identifying a cure
- jmugan 2y agoI believe in search, but it only works if you have an appropriate search space. Chess has a well-defined space but the everyday world does not. The trick is enabling an algorithm to learn its own search space through active exploration and reading about our world. I'm working on that.
- jhawleypeters 2y agoOh nice! The one thing that confused me about this article was what search space the author envisioned adding to language models.
- kragen 2y agothat's interesting; are you building a sort of 'digital twin' of the world it's explored, so that it can dream about exploring it in ways that are too slow or dangerous to explore in reality?
- jmugan 2y agoThe goal is to enable it to model the world at different levels of abstraction based on the question it wants to answer. You can model car as an object that travels fast and carries people, or you can model it down to the level of engine parts. The system should be able to pick the level of abstraction and put the right model together based on its goals.
- kragen 2y agoso then you can search over configurations of engine parts to figure out how to rebuild the engine? i may be misunderstanding what you're doing
- jmugan 2y agoYeah, you could. Or you could search for shapes of different parts that would maximize the engine efficiency. The goal is to simultaneously build a representation space and a simulator so that anything that could be represented could be simulated.
- jhawleypeters 2y agoI think I understand the game space that Leela and now Stockfish search. I don't understand whether the author envisions LLMs searching possibility spaces of 1) written words, 2) models of math / RL / materials science, 3) some smaller, formalized space like the game space of chess, all of the above, or something else. Did I miss where that was clarified?
- fspeech 2y agoHe wants the search algorithm to be able to search for better search algorithms, i.e. self-improving. That would eliminate some of the narrower domains.
- TheRoque 2y agoThe whole premise of this article is to compare the chess state of the art of 2019 with today, and then they start to talk about llms. But chess is a board with 64 squares and 32 pieces, it's literally nothing compared to the real physical world. So I don't get how this is relevant
- dgoodell 2y agoThat’s a good point. Imagine if an LLM could only read, speak, and hear at the same speed as a human. How long would training a model take? We can make them read digital media really quickly, but we can’t really accelerate its interactions with the physical world.
- stephc_int13 2y agoThe author is making a few leap of faith in this article. First, his example of the efficiency of ML+search for playing Chess is interesting but not a proof that this strategy would be applicable or efficient in the general domain. Second, he is implying that some next iteration of ChatGPT will reach AGI level, given enough scale and money. This should be considered hypothetical until proven. Overall, he should be more scientific and prudent.
- aidan_mclau 2y ago[flagged]
- 6510 2y agoI've recently matured to the point where all applications are made of 2 things, search and security. The rest is just things added on top. If you cant find it it isn't worth having.
- brcmthrowaway 2y agoThis strikes me as Lesswrong style pontificating.
- dzonga 2y agoslight step aside - do people at notion realize, their own custom keyboard shortcuts break the habits built on the web. cmd + p -- bring up their own custom dialog. instead of printing the page as one would expect
- sherburt3 2y agoIn VsCode cmd+p pulls up the file search dialog, I don’t think it’s that crazy.
- bob1029 2y agoIt seems there is a fundamental information theory aspect to this that would probably save us all a lot of trouble if we would just embrace it. The #1 canary for me: Why does training an LLM require so much data that we are concerned we might run out of it? The clear lack of generalization and/or internal world modeling is what is really in the way of a self-bootstrapping AGI/ASI. You can certainly try to emulate a world model with clever prompting (here's what you did last, heres your objective, etc.), but this seems seriously deficient to me based upon my testing so far.
- sdenton4 2y agoIn my experience, LLMs do a very poor job of generalizing. I have also seen self supervised transformer methods usually fail to generalize in my domain (which includes a lot of diversity and domain shifts). For human language, you can paper over failure to generalize by shoveling in more data. In other domains, that may not be an option.
- therobots927 2y agoIt’s exactly what you would expect from what an LLM is. It predicts the next word in a sequence very well. Is that how our brains, or even a bird’s brain, for that matter, approach cognition? I don’t think that’s how any animals brain works at all, but that’s just my opinion. A lot of this discussion is speculation. We might as well all wait and see if AGI shows up. I’m not holding my breath.
- stevenhuang 2y agoMost of this is not speculation. It's informed from current leading theories in neuroscience of how our brain is thought to function. See predictive coding and the free energy principle, which states the brain continually models reality and tries to minimize the prediction error. https://en.m.wikipedia.org/wiki/Predictive_coding https://en.m.wikipedia.org/wiki/Predictive_coding
- 2y ago
- zucker42 2y agoIf I had to bet money on it, researchers at top labs have already tried applying search to existing models. The idea to do so is pretty obvious. I don't think it's the one key insight to achieve AGI as the author claims.
- itissid 2y agoThe problem is the transitive closure of chess move is a chess move. The transitive closure of human knowledge and theories to do X is new theories never seen before and no Value function can do that, unless you are also implying theorem proving is included for correctness verification which is also a very difficult search and computationally expensive problem on its own. Also, I think this is instead time to sit back and think what exactly is the thing we value in society as well: Personal(Human) self-sufficiency(I also like to compare this AI to UBI) and thus achievement, which only means Human-in-Loop AI that can help us achieve that and that is specific to each individual, i.e. multi-atttribute value functions whose weights are learned and they change over time. Writing about AGI and defining it to do the "best" search while not talking about what we want it to do *for us* is exactly wrong-headed for these reasons.
- skybrian 2y agoThe article seems rather hand-wavy and over-confident about predicting the future, but it seems worth trying. "Search" is a generalization of "generate and test" and rejection sampling. It's classic AI. Back before the dot-com era, I took an intro to AI course and we learned about writing programs to do searches in Prolog. The speed depends on how long it takes to generate a candidate, how long it takes to test it, and how many candidates you need to try. If they are slow, it will be slow. An example of "human in the loop" rejection sampling is when you use an image generator and keep trying different prompts until you get an image you like. But the loop is slow due to how long it takes to generate a new image. If image generation were so fast that it worked like Google Image search, then we'd really have something. Theorem proving and program fuzzing seem like good candidates for combining search with LLM's, due to automated, fast, good evaluation functions. And it looks like Google has released a fuzzer [1] that can be connected to whichever LLM's you like. Has anyone tried it? [1] https://github.com/google/oss-fuzz-gen https://github.com/google/oss-fuzz-gen
- PartiallyTyped 2y agoBuilding onto this comment; Terrence Tao, the famous mathematician and big proponent of computer aided theorem proving believes ML will open new avenues in the realm of theorem provers.
- sgt101 2y agoSure, but there are grounded metrics there (the theorem is proved, not proved) that allow feedback. Same for games, almost the same for domains with cheap, approximate evaluators like protein folding (finding the structure is difficult, verifying it quite well is cheap). For discovery and reasoning??? Not too sure.
- PartiallyTyped 2y agoLength of proof perhaps?
- YeGoblynQueenne 2y ago
- spencerchubb 2y agoThe branching factor for chess is about 35. For token generation, the branching factor depends on the tokenizer, but 32,000 is a common number. Will search be as effective for LLMs when there are so many more possible branches?
- sdenton4 2y agoYou can pretty reasonably prune the tree by a factor of 1000... I think the problem that others have brought up - difficulty of the value function - is the more salient problem.
- bashfulpup 2y agoThe biggest issue the author does not seem aware of is how much compute is required for this. This article is the equivalent of saying that a monkey given time will write Shakespeare. Of course it's correct, but the search space is intractable. And you would never find your answer in that mess even if it did solve it. I've been building branching and evolving type llm systems for well over a year now full time. I have built multiple "search" or "exploring" algorithms. The issue is that after multiple steps, your original agent, who was tasked with researching or doing biology, is now talking about battleships (an actual example from my previous work). Single step is the only real situation search functions work. Mutli step agents explode to infinite possibilities very very quickly. Single step has its own issues, though. While a zero shot question run 1000 times (eg, solve this code problem), may help find a better solution it's a limited search space (which is a good thing) I recently ran a test of 10k inferences of a single input prompt on multiple llm models varying the input configurations. What you find is that an individual prompt does not have infinite response possibilities. It's limited. This is why they can actually function as llms now. Agents not working is an example of this problem. While a single step search space is massive, it's exponential every step the agent takes. I'm building tools and systems around solving this problem, and to me, a massive search is as far off as saying all we need 100x AI model sizes to solve it. Autonomy =/ (Intelligence or reasoning)
- sorobahn 2y agoI feel like this is a really hard problem to solve generally and there are smart researchers like Yann LeCun trying to figure out the role of search in creating AGI. Yann's current bet seems to be on Joint Embedding Predictive Architectures (JEPA) for representation learning to eventually build a solid world model where the agent can test theories by trying different actions (aka search). I think this paper [0] does a good job in laying out his potential vision, but it is all ofc harder than just search + transformers. There is an assumption that language is good enough at representing our world for these agents to effectively search over and come up with novel & useful ideas. Feels like an open question but: What do these LLMs know? Do they know things? Researchers need to find out! If current LLMs' can simulate a rich enough world model, search can actually be useful but if they're faking it, then we're just searching over unreliable beliefs. This is why video is so important since humans are proof we can extract a useful world model from a sequence of images. The thing about language and chess is that the action space is effectively discrete so training generative models that reconstruct the entire input for the loss calculation is tractable. As soon as we move to video, we need transformers to scale over continuous distributions making it much harder to build a useful predictive world model. [0]: https://arxiv.org/abs/2306.02572 https://arxiv.org/abs/2306.02572
- therobots927 2y ago“Do they know things?” The answer to this is yes but they also think they know things that are completely false. If it’s one thing I’ve observed about LLMs it’s that they do not handle logic well, or math for that matter. They will enthusiastically provide blatantly false information instead of the preferable “I don’t know”. I highly doubt this was a design choice.
- sangnoir 2y ago> “Do they know things?” The answer to this is yes but they also think they know things that are completely false Thought experiment: should a machine with those structural faults be allowed to bootstrap itself towards greater capabilities on that shaky foundation? What would the impact of a near-human/superhuman intelligence that has occasional psychotic breaks it is oblivious of? I'm critical of the idea of super-intelligence bootstrapping off LLMs (or even LLMs with search) - I figure the odds of another AI winter are much higher than those of achieving AGI in the next decade.
- sashank_1509 2y agoI wouldn’t read too much into stockfish beating Leela Chess Zero. My calculator beats GPT-4 in matrix multiplication, doesn’t mean we need to do what my calculator does in GPT-4 to make it smarter. Stockfish evaluates 70 million moves per second (or something in that ballpark). Chess is not such a complicated game that you aren’t guaranteed to find the best move when you evaluate 70 million moves. It’s why when there was an argument whether alpha zero really beat stockfish convincingly in Google’s PR Stunt, a notable chess master quipped “Even god would not be able to beat stockfish this frequently.” , similarly god with all this magical powers would not beat my calculator at multiplication. It says more about the task than about the nature of intelligence.
- Veedrac 2y agoPeople vastly underestimate god. Players aren't just trying not to blunder, they're trying to steer towards advantageous positions. Stockfish could play perfectly against itself every move 100 games in a row, in the classical sense of perfectly, as not in any move blundering the draw, and still be reliably exploited by an oracle.
- galaxyLogic 2y agoHow would search + LLMs work together in practice? How about using search to derive facts from ontological models, and then writing out the discovered facts in English. Then train the LLM on those English statements. Currently LLMs are trained on texts found on the internet mostly (only?). But information on the internet is often false and unreliable. If instead we would have logically sound statements by the billions derived from ontological world-models then that might improve the performance of LLMs significantly. Is something like this what the article or others are proposing? Give the LLM the facts, and the derived facts. Prioritize texts and statements we know and trust to be true. And even though we can't write out too many true statements ourselves, a system that generated them by the billions by inference could.
- scottmas 2y agoBefore an LLM discovers a cure for cancer, I propose we first let it solve the more tractable problem of discovering the “God Cheesecake” - the cheesecake do delicious that a panel of 100 impartial chefs judges to be the most delicious they have ever tasted. All the LLM has to do is intelligently search through the much more combinatorially bounded “cheesecake space” until it finds this maximally delicious cheesecake recipe. But wait… An LLM can’t bake cheesecakes, nor if it could would it be able to evaluate their deliciousness. Until AI can solve the “God Cheesecake” problem, I propose we all just calm down a bit about AGI
- bongodongobob 2y agoYou don't even need AI for that. Try a bunch of different recipes and iterate on it. I don't know what point you're trying to make.
- spencerchubb 2y agoTikTok is the digital version of this
- dontreact 2y agoThese cookies were very good, not God level. With a bit of investment and more modern techniques I think you could make quite a good recipe, perhaps doing better than any human. I think AI could make a recipe that wins in a very competitive bake-off, but it’s not possible or for anyone to win with all 100 judges. https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/46507.pdf https://static.googleusercontent.com/media/research.google.c...
- IncreasePosts 2y agoHeck, even theoretically 100% within the limitations of an LLM executing on a computer, it would be world changing if LLMs could write a really, really good short story or even good advertising copy.
- eru 2y agoThey are getting better and better. I am fairly sure the short stories and advertising copy you can produce by pushing current techniques harder will also improve. I don't know whether current techniques will be enough to 'write a really, really good short story', but I'm willing to bet we'll get there soon enough (whether that'll involve new techniques or not).
- omneity 2y agoThe post starts with a fascinating premise, but then falls short as it does not define search in the context of LLMs, nor does it explain how “Pfizer can access GPT-8 capabilities today with more inference compute”. I found it hard to follow and I am an AI practitioner. Could someone please explain more what could the OP mean? To me it seems that the flavor of search in the context of chess engines (look several moves ahead) is possible precisely because there’s an objective function that can be used to rank results, i.e. which potential move is “better” and this is more often than not a unique characteristic of reinforcement learning. Is there even such a metric for LLMs?
- sgt101 2y agoyeah - I think that that's waht they mean and I think that there isn't such a metric. I think people will try to do adversarial evaluation but my guess is that it will just tend to the mean prediction. The other thing is that LLM inference isn't cheap. The trade off between inference costs and training costs seems to be very application specific. I suppose that there are domains where accepting 100x or 1000x inference costs vs 10x training costs makes sense, maybe?
- qnleigh 2y agoThank you, I am also very confused on this point. I hope someone else can clarify. As a guess, could it mean that you would run the model forward a few tokens for each of its top predicted tokens, keep track of which branch is performing best against the training data, and then use that information somehow in training? But search is supposed to make things more efficient at inference time and this thought doesn't do that...
- deleted 2y ago[deleted]
- amandasystems 2y agoThis feels a lot like generation 3 AI throwing out all the insights from gens 1 and 2 and then rediscovering them from first principles, but it’s difficult to tell what this text is really about because it lumps together a lot of things into “search” without fully describing what that means more formally.
- PontifexMinimus 2y agoIndeed. It's obvious what search means for a chess program -- its the future positions it looks at. But it's less obvious to me what it means for an LLM.
- YeGoblynQueenne 2y ago>> She was called Leela Chess Zero — ’zero’ because she started knowing only the rules. That's a common framing but it's wrong. Leela -and all its friends- have another piece of chess-specific knowledge that is indispensable to their performance: they have a representation of the game of chess -a game-world model- as a game tree, divided in plys: one ply for each player's turn. That game tree is what is searched by adversarial search algorithms, such as minimax or Monte Carlo Tree Search (MCTS; the choice of Leela, IIUC). More precisely modelling a game as a game tree applies to many games, not just chess, but the specific brand of game tree used in chess engines applies to chess and similar, two-person, zero-sum, complete information board games. I do like my jargon! For other kinds of games, different models, and different search algorithms are needed, e.g. see Poker and Libratus [1]. The need for such a game tree, such a model of a game world, is currently impossible to go without, if the target is superior performance. The article mentions no-search algorithms and briefly touches upon their main limitation (i.e. "why?"). All that btw is my problem with the Bitter Lesson: it is conveniently selective with what it considers domain knowledge (i.e. a "model" in the sense of a theory). As others have noted, e.g. Rodney Brooks [2], Convolutional Neural Nets have dominated image classification thanks to the use of convolutional layers to establish positional invariance. That's a model of machine vision invented by a human, alright, just as a game-tree is a model of a game invented by a human, and everything else anyone has ever done in AI and machine learning is the same: a human comes up with a model, of a world, of an environment, of a domain, of a process, then a computer calculates using that model, and sometimes even outperforms humans (as in chess, Go, and friends) or at the very least achieves results that humans cannot match with hand-crafted solutions. That is a lesson to learn (with all due respect to Rich Sutton). Human model + machine computation has solved every hard problem in AI in the last 80 years. And we have no idea how to do anything even slightly different. ____________________ [1] https://en.wikipedia.org/wiki/Libratus https://en.wikipedia.org/wiki/Libratus [2] https://rodneybrooks.com/a-better-lesson/ https://rodneybrooks.com/a-better-lesson/
- nojvek 2y agoWe haven’t seen algorithms that build world models by observing. We’ve seen hints of it but nothing human like. It will come eventually. We live in exciting times.
- schlipity 2y agoI don't run javascript by default using NoScript, and something amusing happened on this website because of it. The link for the site points to a notion.site address, but attempting to go to this address without javascript enabled (for that domain) forces a redirect to a notion.so domain. Attempting to visit just the basic notion.site address also does this same redirection. What this ends up causing is that I don't have an easy way to use NoScript to temporarily turn on javascript for the notion.site domain, because it never loads. So much for reading this article.
- ajnin 2y agoOT but this website completely breaks arrow and page up/down scrolling, as well as alt=arrow navigation. Only mouse scrolling works for me (I'm using Firefox). Can't websites stop messing with basic browser functionality for no valid reason at all ?
- awinter-py 2y agojust came here to upvote the alphago / MCTS comments
- kunalgupta 2y agoThis is one of my favorite reads in a while
- Hugsun 2y agoA big problem with the conclusions of this article is the assumptions around possible extrapolations. We don't know if a meaningfully superintelligent entity can exist. We don't understand the ingredients of intelligence that well, and it's hard to say how far the quality of these ingredients can be improved, to improve intelligence. For example, an entity with perfect pattern recognition ability, might be superintelligent, or just a little smarter than Terrance Tao. We don't know how useful it is to be better at pattern recognition to an arbitrary degree. A common theory is that the ability modeling processes, like the behavior of the external world is indicative of intelligence. I think it's true. We also don't know the limitations of this modeling. We can simulate the world in our minds to a degree. The abstractions we use make the simulation more efficient, but less accurate. By this theory, to be superintelligent, an entity would have to simulate the world faster with similar accuracy, and/or use more accurate abstractions. We don't know how much more accurate they can be per unit of computation. Maybe you have to quadruple the complexity of the abstraction, to double the accuracy of the computation, and human minds use a decent compromise that is infeasible to improve by a large margin. Maybe generating human level ideas faster isn't going to help because we are limited by experimental data, not by the ideas we can generate from it. We can't safely assume that any of this can be improved to an arbitrary degree. We also don't know if AI research would benefit much from smarter AI researchers. Compute has seemed to be the limiting factor at almost all points up to now. So the superintelligence would have to help us improve compute faster than we can. It might, but it also might not. This article reminds me of the ideas around the singularity, by placing too much weight on the belief that any trendline can be extended forever. It is otherwise pretty interesting, and I'm excitedly watching the 'LLM + search' space.
- TZubiri 2y ago"In 2019, a team of researchers built a cracked chess computer. She was called Leela Chess Zero — ’zero’ because she started knowing only the rules. She learned by playing against herself billions of times" This is a gross historical misunderstanding or misrepresentation. Google accomplished this feat. Then the OS+academics reversed engineered/duplicated the study
- RA_Fisher 2y agoLearning tech improves on search tech (and makes it obsolete), because search is about distance minimization not integration of information.