6 ms·
Language models as compilers: Simulating pseudocode execution
- Voultapher 2y agoNon deterministic compilers, yay! Where do I sign up? In more seriousness, miscompilations or in general unexpected behavior caused by layers below you are expensive to find and fix. I think LLMs have a long way to go before such use cases seem appealing to me.
- ginko 2y agoEven regular compilers need quite a bit of nudging to give deterministic results.
- Phillipharryt 2y agoCorrect me if I'm wrong here, but I am under the impression they're only non-deterministic in the practical sense (i.e, it produces this output on my machine, I can't know what minute differences there are on your machine), but that's not non-deterministic in the truest sense. If you have completely identical inputs you will get the exact same output, ergo, deterministic.
- layer8 2y agoYou are correct. Compilers are deterministic, but reproducible builds can be a challenge.
- deleted 2y ago[deleted]
- tiborsaas 2y agoIt's better to have a non-deterministic compiler for a task that would be really hard to write an algorithm for otherwise.
- clbrmbr 2y agoBut are LLMs really better at algorithm writing than you? I’ve found that they work best when I’ve already pseudocoded the algorithm.
- tiborsaas 2y agoI didn't mean making it write the actual code and using that, but there are tasks that are more error prone or near impossible to write with a traditional approach. So using some zero shot prompting is better than running code. That's why I find non-determinism as acceptable when otherwise it would be a pain to do something similar.
- taneq 2y agoNon-determinism is an implementation detail, not an intrinsic property, as I understand it (at least as long as you're setting temperature to zero). I likewise don't really think LLMs are the right tool for this job, though. There's a whole class of systems that we built because humans take a long time to learn new skills, are fallible and non-repeatable, and get bored easily. Compilers are in this group along with sewing machines, CNC machines, automatic gearboxes, and design rules checking in CAD. Maybe they could provide heuristics for optimising compilers with the output run through a formal verification check afterwards?
- knightoffaith 2y ago>Non-determinism is an implementation detail, not an intrinsic property, as I understand it (at least as long as you're setting temperature to zero). Right. A transformer outputs a probability distribution over all possible tokens from which the next token is sampled and then appended to the input sequence, at which point the process repeats. Temperature controls the entropy of the distribution - higher temperature, higher entropy, conversely, lower temperature, lower temperature. Technically zero temperature involves dividing by zero, so under the hood it's simply set to be an epsilon so small that the entropy of the distribution is low enough that sampling from it always effectively gives one token - the token with the highest probability. And so at every step in inference, the highest probability token is emitted.
- deleted 2y ago[deleted]
- eeue56 2y agoI (kinda) solved this with neuro-lingo[0] with the concept of pinning. Basically, once you have a version of a function implementation that works, you can pin it and it won't be regenerated when it's "compiled". The alternative approach would be to have tests be the only code a developer writes, and then make LLMs generate code to match the implementation for those, running the tests to ensure it's valid. - [0] https://github.com/eeue56/neuro-lingo https://github.com/eeue56/neuro-lingo
- lionkor 2y agoHow is it better than a compiler written by people?
- eeue56 2y agoI wrote a toy language along these lines a while back[0]. Basically, types and function signatures, with comments in English, produce a valid program. You write a type and a comment, and the compiler goes through GPT to run the code. Fun novel idea. [0] - https://github.com/eeue56/neuro-lingo https://github.com/eeue56/neuro-lingo
- Mathnerd314 2y agoThis paper suggests that pseudocode is more understandable to LLM's than natural language. So building off of that, one should not write comments in the method bodies, one should translate a function body (in mixed English and pseudocode) as a unit. LLM's can memorize a lot but there is a limit so I guess it would have to be structured as a multi-level translation. You mention in the Readme that it is persnickety, this is probably due to the use of "auto-complete" in the prompt. I would use an intro like "here is a function body in pseudocode", then prompts "split it into chunks by inserting labels", "translate whole body preserving chunk labels", "re-translate each chunk using whole body as context", "fix type errors". I think the biggest issues besides prompt engineering are cost and latency though.
- novideogame 2y agoI think the title is a little misleading. The main difference between this paper and CoC (Chain of Code) is that the LLM is instructed to make a plan to solve all the given instances and then code that plan in pseudocode, while in CoC the plan is to solve the single instance given. From the paper: The main difference between THINK-AND-EXECUTE and CoC is that we use pseu- docodes which are generated to express logic shared among the tasks instances, while CoC incorporates pseudocode as part of the intermediate reasoning steps towards the solution of a given instance. Hence, the results indicate the advantages of applying pseudocode for the generation of task-level instruction over solely using them as a part of rationales. I find the phrase "as a part of rationales" a little strange, but English is not my native language.
- ingigauti 2y agoCouple of weeks ago I published a new programming language called Plang (as in pseudo language) that uses LLM to translate user intent into executable code, basically LLM as a compiler. It saves you incredible amount of work, cutting code writing down by 90%+. The built code is deterministic(it will never change after build) and as a programmer you can validate the code that will be executed. It compiles to C#, so it handles GC, encoding, etc. that languages need to solve, so I can focus on other areas. Plang also has some features that other language don't have, e.g. events on variables, built in identity and interesting(I think) approach to privacy. I have not been advertising to much since it is still early development and I create still to many breaking changes, but help is welcome(and needed) so if it something that is interesting to you the repo is at https://github.com/plangHQ https://github.com/plangHQ
- layer8 2y agoIn my experience it’s not exactly trivial to validate code you didn’t write yourself, because you have to think through it in similar depth to when you write it yourself. While on the one hand you save the time of coming up with a solution, the task of merely verifying an existing solution is also more tedious, because it isn’t intermixed with the problem-solving activity that you perform when writing the code yourself. There is an increased risk to fall for code that looks correct at first blush but is still subtly wrong, because you didn’t spend time iterating on it. It doesn’t seem plausible to me that you would save 90% of the work, unless it’s boiler plate-heavy code that requires only little analytical though.
- ingigauti 2y agoI agree with you. I never liked how AI is really generating ton of code for us, then you need to read through it and understand it. Plus, the fail rate is to high. That is why I design the language the way it is. You must define each step you want to happen in your application. Lets take for example user registration, it looks like this --- plang code --- CreateUser - Make sure %password% and %email% is not empty - Hash %password%, write to %hashedPassword% - Insert into users, %hashedPassword%, %email% - Post, create user in MailChimp, Bearer:%Settings.MailChimpApi%, %email% - Create bearer token from %email%, write to %bearer% - Write %bearer% to web response --- plang code --- That is an executable code in plang. It's easy to read through and understand. You need to have domain knowledge, such as what is hashing and bearer token. You are still programming, just at higher level. Validating what will execute, you need to learn, just like with any language, but it is relatively simple and you start to trust the result with time(at least I have) Compared to the 130 lines or so of code in C# for the same logic, https://gist.github.com/ingig/491ac9b13d65f40cc24ee5aed0408be3 https://gist.github.com/ingig/491ac9b13d65f40cc24ee5aed0408b... That´s about 95% reduction of code, and I see this repeatedly.
- imranq 2y agoReading the paper, the connection to compilers is more of an analogy rather than a direct technical link. The authors propose using an LLM to reframe the task as high level psuedocode, and then reason on that code on the specific details of the task No compilers were used or compiled - no real code was generated or executed. Its just the idea that a programming language syntax has good structure to process details, and a way to interpret some of the results. Many of the other comments here seem like they didn't read the paper at all and are reacting to the headline
- 29athrowaway 2y agoUp next: A LLM that can tell me if a program stops
- emmender2 2y agoResearchers are trying their damndest to build a "reasoning" layer using LLMs as the foundation. But, they need to go back to the drawing-board and understand from first principles what it means to reason. For this in my view, they need to go back to epistemology (and refer to Peirce and logicians like him).
- spxneo 2y agoThis seems quite promising. Using pseudo-code as an intermediary step isn't new but seems like this takes it a bit further. Will need to see some code and test it out.
- Mathnerd314 2y agoThe phase 2 prompt is complete, but the phase 3 prompt's initial part ends in "When constructing the main function, ...", and no mention of random seeds, so I guess this paper is not reproducible at all.
- inciampati 2y agoIt's going to be really fascinating to see this applied instead of chain of thought and other kinds of reasoning approaches, because it's generic. It should in principle work on every kind of LLM.
- pkoird 2y agoAny sufficiently advanced LLM is indistinguishable from Prolog. I half-jest but I envision the direction of LLM research to head towards a parser-oriented setup where LLMs merely extract the entities and relations and the actual logic is done by a logical engine such as Prolog.
- westoncb 2y agoI think this is an important perspective. It helps clarify prompting as well because, used in a certain way, they are effectively natural language specification of constraints which have the effect of 'partially configuring' the network so that its inference selects from the configuration space of its still free params For example, if you tell it to reply in JSON (and it obeys), you've just constrained its search space in a particular way. There is space for very interesting informal programming that can be done from this perspective, setting up constraints and then allowing inference to solve within them. I've been using this heavily. When I was first getting deep into LLM stuff a few months ago and contemplating latent space my main characterization was that much of its high level behavior can be usefully grappled with by viewing it as a kind of 'learned geometric prolog'. I did a bunch of illustrations and talked about some of these ideas here if anyone's curious: https://x.com/Westoncb/status/1757910205478703277 https://x.com/Westoncb/status/1757910205478703277 (I think I mostly dropped the prolog terminology in that presentation because not everyone knows about it)
- thesz 2y agohttps://en.wikipedia.org/wiki/Cyc#MathCraft https://en.wikipedia.org/wiki/Cyc#MathCraft Quote: One Cyc application aims to help students doing math at a 6th grade level, helping them much more deeply understand that subject matter... Unlike almost all other educational software, where the computer plays the role of the teacher, this application of Cyc, called MathCraft, has Cyc play the role of a fellow student who is always slightly more confused than you, the user, are about the subject. This is from 2017. I haven't seen anything like this using LMs in 2017 and I suspect it is still hard for LLMs today. Cyc is the huge reasoning engine. You can call it Prolog, if you want. I won't.
- 2y ago
- skeledrew 2y agoSeeing this makes me want to reactivate an old project[0]. Been thinking more and more that LLMs could give it superpowers. [0] https://pypi.org/project/neulang/ https://pypi.org/project/neulang/
- m3kw9 2y agoIf you train a LLM to compile, you probably also want to set the randomness to zero, if that is the case you’ve just “brute forced” an actual compiler
- danielmarkbruce 2y agoyou don't want a creative compiler?
- fxcao 2y agoIn this specific use case I think we should avoid any creative aspect in the behavior of the LLM. Compiling might look like a "word-for-word" translation in a certain manner. Isn't it ?
- layer8 2y agoI see no inherent conflict between creativity and determinism. You could give the compiler a seed value to perturbate the solution space.
- m3kw9 2y agoWhy vote down a rhetorical question?
- jumploops 2y agoEnglish is terribly imprecise, so it makes sense to use pseudo instructions to improve the bounds/outcome of a language model’s execution. I do wonder how long hacks like this will be necessary; as it stands, many of these prompting techniques are essentially artificially expanding the input to enhance reasoning ability (increasing tokens, thus increasing chance of success).
- nimbleal 2y agoI hope there are more models trained on more precise inputs going forward. I understand that natural language feels the most futuristic but while it has the lowest barrier to entry it’s not only imprecise but also slow. Visual approaches (for example control nets in stable diffusion, image as input in Chat GPT, though both of these are somewhat bolted on), 2D semi-natural languages all merit further inquiry. Another (and perhaps the ultimate) possibility is to have some way —- perhaps through simulations —- to directly expose the model to the problem, rather than having a human/natural language intermediary.