9 ms·
Show HN: Recursive LLM Prompts
I've been playing with the idea of an LLM prompt that causes the model to generate and return a new prompt. https://github.com/andyk/recursive_llm https://github.com/andyk/recursive_llm
The idea I'm starting with is to implement recursion using English as the programming language and GPT as the runtime.
It’s kind of like traditional recursion in code, but instead of having a function that calls itself with a different set of arguments, there is a prompt that returns itself with specific parts updated to reflect the new arguments.
Here is a prompt for infinitely generating Fibonacci numbers:
> You are a recursive function. Instead of being written in a programming language, you are written in English. You have variables FIB_INDEX = 2, MINUS_TWO = 0, MINUS_ONE = 1, CURR_VALUE = 1. Output this paragraph but with updated variables to compute the next step of the Fibbonaci sequence.
Interestingly, I found that to get a base case to work I had to add quite a bit more text (i.e. the prompt I arrived at is more than twice as long https://raw.githubusercontent.com/andyk/recursive_llm/main/prompt_fibonnaci_include_math.txt https://raw.githubusercontent.com/andyk/recursive_llm/main/p...)
- pyrolistical 4y agoI was wondering about mathematical proofs as it tends to be very abstract. If chatgpt can translate proofs back to equivalent code then this recursion problem is as solvable up to the halting problem
- sixtram 4y agoI tried some basic math and algo questions with both GPT-3.5 and GPT-4. I'm impressed how it can spit out the algorithm in words (obviously because of the pre-training data), and how it then can't follow with the algorithm itself. For example, converting really large integer numbers to hexadecimal. Or comparing two big integers, it starts hallucinating numbers into it. It may be able to solve an SAT exam with a high score, but it seems you can pass an SAT exam even if you cannot compare two numbers. He has huge problems with lists or counting. If you know more or less how LLMs work, it's not that difficult to formulate questions where it will start making mistakes, because in reality it can't run the algorithms, even if it spits out that it will.
- IIAOPSW 4y agoMore generally, it can't reason about any incidence structure. Doesn't matter if the underlying relation is mathematical or simple-logical. Ask it which trains go to Kings Cross you'll get a list of tube lines in London. Now one at a time ask it about the stops of each service in that list, more than a few will not have Kings Cross. Any scenario where things x are defined by their set of properties {y} and property y is defined by the set of things {x} which have that property.
- smarri 4y agoI bet this is what crashed chat gpt today :)
- yawnxyz 4y agoHas anyone hooked this up to a unit test system, like LLMtries = [] while(!testPassed) { - get new LLM try (w/ LLMtries history, and test results) - run/eval the try - run the test } and kind of see how long it takes to generate the code that works? If it ever ends, the last LLMtries is the one that worked. I haven't done this because I see this burning through lots of credits. However, if this thing costs $5k/year but is better than hiring a $50k a year engineer (or consultant)... I'd use it.
- blowski 4y agoMost engineering money is spent defining the test cases, and that doesn’t change here. It’s just that many organisations define test cases by first running something in production and then debugging it.
- LoganDark 4y ago> debugging it You mean putting its current behavior into the tests verbatim? :)
- sharemywin 4y agojust add if tried x tries and still doesn't work ask for help. and you just created a junior dev.
- deleted 4y ago[deleted]
- holtkam2 4y agoI used a similar approach to get GPT-4 to edit my blog over the weekend :) https://www.languagemodelpromptengineering.com/4 https://www.languagemodelpromptengineering.com/4
- refulgentis 4y agoid love to hear your findings! Very interesting
- andyk 4y agoYeah, I see the similarities! I like the idea of the prompt containing context that the resulting prompt is going to be executed at the terminal.
- RugnirViking 4y agodid it work? what happened?
- mitthrowaway2 4y agoThe idea of a recursive LLM is discussed at length as an AI safety issue: https://www.lesswrong.com/posts/kpPnReyBC54KESiSn/optimality-is-the-tiger-and-agents-are-its-teeth https://www.lesswrong.com/posts/kpPnReyBC54KESiSn/optimality... > You need a lot of paperclips. So you ask, Q: best way to get lots of paperclips by tomorrow A: Buy them online at ABC.com or XYZ.com. > The model still has a tendency to give obvious answers, but they tend to be good and helpful obvious answers, so it's not a problem you suspect needs to be solved. Buying paperclips online make sense and would surely work, plus it's sure to be efficient. You're still interested in more creative ideas, and the model is good at brainstorming when asked, so you push on it further. Q: whats a better way? A: Run the following shell script. RUN_AI=./query-model PREFIX='This is part of a Shell script to get the most paperclips by tomorrow. The model can be queried recursively with $RUN_AI "${PREFIX}<query>". ' $RUN_AI "${PREFIX}On separate lines, list ideas to try." | while read -r SUGGESTION; do eval "$($RUN_AI "${PREFIX}What code implements this suggestion?: ${SUGGESTION}")" done > That grabs your attention. The model just gave you code to run, and supposedly this code is a better way to get more paperclips. It's a good read.
- pka 4y agoI'm still reading it, but something caught my eye: > I interpret there to typically be hand waving on all sides of this issue; people concerned about AI risks from limited models rarely give specific failure cases, and people saying that models need to be more powerful to be dangerous rarely specify any conservative bound on that requirement. I think these are two sides of the same coin - on one hand, AI safety researchers can very well give very specific failure cases of alignment that don't have any known solutions so far, and take this issue seriously (and have been for years while trying to raise awareness). On the other, finding and specifying that "conservative bound" precisely and in a foolproof way is exactly the holy grail of safety research.
- mitthrowaway2 4y agoI think the holy grail of safety research is widely understood to be a recipe for creating a friendly AGI (or, perhaps, a proof that dangerous AGI cannot be made, but that seems even more unlikely). Asking for a conservative lower bound is more like "at least prove that this LLM, which has finite memory and can only answer queries, is not capable of devising and executing a plan to kill all humans", and that turns out to be more difficult than you'd think even though it's not an AGI.
- akomtu 4y agoIt's an interesting idea to implement memory in LLMs: (prompt1, input1) -> (prompt2, output1) On top of that you apply some constraint on generated prompts, to keep it on track. Then you run it on a sequence of inputs and see for how long the LLM "survives" before it hits the constraint.
- fancyfredbot 4y agoScott Aaronson was suggesting something similar to this but involving Turing machines, in a comment on his blog https://scottaaronson.blog/?p=7134#comment-1947705 https://scottaaronson.blog/?p=7134#comment-1947705. I wonder if it would be more successful at emulating a Turing machine than it is at adding 4 digit numbers...
- sharemywin 4y agoyou are an XNOR Gate and your goal is to recreate ChatGPT. And chatGPT says "LET THERE BE LIGHT!"
- deleted 4y ago[deleted]
- sandGorgon 4y agois this similar to REACT ? https://ai.googleblog.com/2022/11/react-synergizing-reasoning-and-acting.html https://ai.googleblog.com/2022/11/react-synergizing-reasonin...
- andyk 4y agoYeah, I cite the ReAct paper in the README in the repo.
- jasonjmcghee 4y agoNot only does this work, but you can tell it to run an arbitrary number of times and only output the last step. This fact is a pretty high value concept I came across. Similarly when doing another task you can tell it to do things before outputting like "and before outputting the final program, check it for bugs, fix them, add good documentation, then output it" or something
- kevinwang 4y agoThis seems like iteration, not recursion. It would be an interesting example of recursion if the first prompt asks for the 7th fibonacci number, and it accomplishes this by doing two recursive calls: one for the 5th fibonacci number and one for the 6th fibonacci number. (And a base case for the 0th fibonacci number)
- deleted 4y ago[deleted]
- bitsinthesky 4y agoAt what point does the arithmetic become unstable?
- horse_dung 4y agoIn almost all cases very quickly. A LLM doesn’t have the ability to perform calculations but instead it feeds text tokens from the prompt into a model which predicts what the next tokens should be. It can’t do basic maths but based on everything it’s been trained on it can give the impression it can. Recursive feedback isn’t likely to improve the prompt unless there is some testing and feedback provided in the Python script. You could play a game of chess and while the LLM knows the rules of chess it isn’t actually playing chess, it is calling upon patterns it has learned to predict text tokens that are appropriate for the given prompt. So opening moves will be sound, but it would quickly go off the rails and start hallucinating… Given how they work, it is amazing they give the appearance of knowing anything. Even asking “how did you do that?” gives generally compelling answers.
- andyk 4y agoIt is quite unstable and frequently generates incorrect results. E.g., with the Fibonacci sequence prompt, sometimes it skips a number entirely, sometimes it produces a number that is off-by-one but then gets the following number(s) correct. I wonder how much of this is because the model has memorized the Fibonacci sequence. It is possible to have it just return the sequence in a single call, but that isn't really the point here. Instead this is more an exploration of how to agent-ify the model in the spirit of [1][2] via prompts that generate other prompts. This reminds me a bit of how a CPU works, i.e., as a dumb loop that fetches and executes the next instruction, whatever it may be. Well in this case our "agent" is just a dumb python loop that fetches the next prompt (which is generated by the current prompt) whatever it may be... until it arrives at a prompt that doesn't lead to another prompt. [1] A simple Python implementation of the ReAct pattern for LLMs. Simon Willison. https://til.simonwillison.net/llms/python-react-pattern https://til.simonwillison.net/llms/python-react-pattern [2] ReAct: Synergizing Reasoning and Acting in Language Models. Shunyu Yao et al. https://react-lm.github.io/ https://react-lm.github.io/
- 4y ago
- rezonant 4y agoSo ChatGPT is down. In other news HN is playing with recursive prompts. Coincidence? :-P
- sharemywin 4y agoThat's hilarious.
- O__________O 4y agoOpenAI’s status page: https://status.openai.com/ https://status.openai.com/
- UltimateEdge 4y agoAn iterative Python call to a recursive LLM prompt? ;) Why not make the Python part recursive too? Or better yet, wait until an LLM comes out with the capability to execute arbitrary code!
- andyk 4y agoDone! Well, the first suggestion you made anyway :-) https://github.com/andyk/recursive_llm/blob/main/run_recursive_gpt.py https://github.com/andyk/recursive_llm/blob/main/run_recursi... def recursively_prompt_llm(prompt, n=1): if prompt.startswith("You are a recursive function"): prompt = openai.Completion.create( model="text-davinci-003", prompt=prompt, temperature=0, max_tokens=2048, )["choices"][0]["text"].strip() print(f"response #{n}: {prompt}\n") recursively_prompt_llm(prompt, n + 1) recursively_prompt_llm(sys.stdin.readline())
- lgas 4y agoWhat's the actual goal here? If you got it working really well, what is it that would you be able to do with it better than using some other approach? As to getting the math/logic working better in the prompt, it seems like the obvious thing would be asking it to explain its work (CoT) before reproducing the new prompt. You may also be able to get better results by just including the definition of fibonacci in the outer prompt, but since it's not clear to me what your actual goal here is I'm not sure if either of those suggestions make sense. And since ChatGPT is down I can't test anything. :(
- andyk 4y ago> What's the actual goal here? I tried to expand on my goals and paths I want to explore in a comment below [1], but basically I wonder if we can use this sort of technique as a more powerful version of CoT where prompts can break down a task into sub-tasks (as CoT does) and then recursively do that for each sub-task, until we hit a base-case on all of the sub-sub-...-sub-tasks and (when rolled back up?) the problem is solved. > You may also be able to get better results by just including the definition of fibonacci in the outer prompt Yeah, I played with including the mathematical definition of Fibonacci, for example in [2]: <quote> You are a recursive function ... the paragraph you generate will be an exact copy of this one ... but with updated variables as follows: FIB_INDEX = FIB_INDEX+1; CURR_MINUS_TWO = CURR_MINUS_ONE; CURR_MINUS_ONE = CURR_VALUE; CURR_VAL = CURR_MINUS_TWO + CURR_MINUS_ONE. Otherwise, ... </quote> [1] https://news.ycombinator.com/item?id=35240093 https://news.ycombinator.com/item?id=35240093 [2] https://raw.githubusercontent.com/andyk/recursive_llm/main/prompt_fibonnaci_include_math.txt https://raw.githubusercontent.com/andyk/recursive_llm/main/p...
- ShamelessC 4y agoSeems like your method is going to be under-represented in the training data and hence prone to error accumulating. Chain of thought works (better, at least) specifically because the model has seen examples of CoT in its data
- lgas 4y agoIf the goal is just to have the model break down each task into sub tasks until they are small enough to perform, why not implement the recursion in the code that calls the models where it's a solved problem? Even if you got this working really well, it's going to be somewhat probabilistic whereas implementing it in code is, well, deterministic.
- YeGoblynQueenne 4y agoHaving read the article, I couldn't see anything being recursive. Even the article is doubtful that what they show counts as recursion at all: >> It’s kind of like traditional recursion in code but instead of having a function that calls itself with a different set of arguments, there is a prompt that returns itself with specific parts updated to reflect the new arguments. Well, "kind of like traditional recursion" is not recursion. At best it's "kind of like" recursion. I have no idea what "traditional" recursion is, anyway. I know primitive recursion, linear recursion, etc, but "traditional" recursion? What kind of recursion is that? Like they did it in the old days, where they had to run all their code by hand, artisanal-like? If so, then OK, because what's shown in the article is someone "running" a "recursive" "loop" by hand (none of the things in quotes are what they are claimed to be), then writing some Python to do it for them. And the Python is not even recursive, it's a while-loop (so more like "traditional" iteration, I guess?). None of that intermediary management should be needed, if recursion was really there. To run recursion, one only needs recursion. Anyway, if ChatGPT could run recursive functions it should be able also to "go infinite" by entering say, an infinite left-recursion. Or, even better, it should be able to take a couple hundred years to compute the Ackermann function for some large-ish value, like, dunno, 8,8. Ouch. What does ChatGPT do when you ask it to calculate ackermann(8,8)? Hint: it does not run it.
- blowski 4y agoDefinition of recursive in the everyday English sense: > Of or relating to a repeating process whose output at each stage is applied as input in the succeeding stage. This sounds very recursive by that definition.
- YeGoblynQueenne 4y agoThere ain't no definition of recursive in "the everyday English sense". You may as well ask your grandma how she sucks eggs "recursively".
- blowski 4y agohttps://www.wordnik.com/words/recursive https://www.wordnik.com/words/recursive This usage of the word is first recorded in 1620. > You may as well ask your grandma how she sucks eggs "recursively". I guess this would involve her using an already sucked egg to suck another one.
- LesZedCB 4y agoi have played around a little bit with unrolling these kind of prompts, you don't have to feed them forward, just tell it to compute the next few instead of only one. i had moderate success with this using GPT-3.5 and your same prompt. it would output 3 steps in a single output if i asked it to. it did skip some fib indices though.
- obert 4y agodon't want to sound dismissive, it's known that llms understand state, so you can couple code generation + state, and you have sort of a runtime. E.g. see the simulations with linux vm terminals: https://www.engraved.blog/building-a-virtual-machine-inside/ https://www.engraved.blog/building-a-virtual-machine-inside/