5 ms·
There are a couple different approaches: - Use multi-shot prompting with something like guardrails to try prompting a commercial model until it works. [1] - U
by newhouseb 3y ago
There are a couple different approaches:
- Use multi-shot prompting with something like guardrails to try prompting a commercial model until it works. [1]
- Use a local model with a final layer that steers token selection towards syntactically valid tokens [2]
[1] https://github.com/ShreyaR/guardrails https://github.com/ShreyaR/guardrails
[2] "Structural Alignment: Modifying Transformers (like GPT) to Follow a JSON Schema" @ https://github.com/newhouseb/clownfish https://github.com/newhouseb/clownfish (full disclosure: this is my work)
- trifurcate 3y agoRegarding [2], dang, I am working on exactly this! I mean, it's not that novel of a technique once you start controlling the sampling process directly, but you beat me to the punch. This technique generalizes to pretty much any grammar one can specify. I weakly hypothesize that by making it impossible for the LM to output syntactically invalid text, the model's task performance improves not just because all of its outputs are valid, but also because part of the model's "processing power" gets "rerouted" from trying to understand and follow the grammar it's writing, to applying improved reasoning overall.
- onesphere 3y agoWhat if the output schema were something like instruction code? Just get rid of the need for programming languages altogether.
- mirker 3y agoThere’s actually a few papers already on constrained decoding. I won’t link them but if you go on arxiv and really look you will find a couple in the past year.
- newhouseb 3y agoNice! I've been wondering similar things about whether you could use this to eek out more intelligence through methods like these, to quote the end of my write up: > Does structured decoding increase the observability of emergent world models in these models? To make an analogy: I may not represent an opinion of how something works if I am not confident in it, but if I am forced to present an opinion we might find out that I in fact have (or have not) grasped something. In practice, however, without tight integration with beam search, the autoregressive nature of these models means that the syntactic steering may result in the models rabbit-holing themselves without forward looking visibility that's obvious from the defined grammar. I.e. if it was forced to choose between "Don't Jump" and "Do run" in some hypothetical example, the set of tokens that it would likely be deciding between is "Don't" and "Do" with no idea what is going to end up syntactically required after those tokens.
- say_it_as_it_is 3y agoSpends years working on an AI solution to problems caused by using postgres as a KV store. That's quite a branch.
- ukuina 3y agoI like that you use a local model to start off; why switch to OpenAI for tokenization?
- newhouseb 3y agoThe code supports both local and OpenAI as a backend, I added OpenAI as a backend because their models are still miles better than anything I can run locally (even 65B LLaMA)