17 ms·
A guidance language for controlling LLMs
- sharemywin 3y agoIt does look like it makes easier to code against a model. But, is this supposed to work along side lang-chain or hugging face agents or as an alternative to?
- evanmays 3y agoIt's in langchain competitor territory but also much lower level and less opinionated. I.e. Guidance has no vector store support but it does manage caching Key/Value on the GPU which can be a big latency win
- ttul 3y agoThe first commit was on November 6th, but it didn't show up in Web Archive until May 6th, suggesting it was developed mostly in private and in parallel with LangChain (LangChain's first commit in Github is about October 24th). Microsoft's code is very tidy and organized. I wonder if they used this tool internally to support their LLM research efforts.
- Terretta 3y agoSomething like this could be a helpful framework to mock and research-iterate purpose-directed tools such as Microsoft GitHub's CoPilot for VSCode.
- slundberg 3y agoAs others mentioned, this was initially developed before LangChain became widely used. Since it is lower level, you can leverage other tools, like any vector store interface you like such as in LangChain. Writing complex chain of thought structure is much more concise in guidance I think since it tries to keep you as close to the real strings going into the model as possible.
- ntonozzi 3y agoHow does this work? I've seen a cool project about forcing Llama to output valid JSON: https://twitter.com/GrantSlatton/status/1657559506069463040 https://twitter.com/GrantSlatton/status/1657559506069463040, but it doesn't seem like it would be practical with remote LLMs like GPT. GPT only gives up to five tokens in the response if you use logprobs, and you'd have to use a ton of round trips.
- joshka 3y agoYeah, I'm also curious about a) round trips and b) how much would have to be doubled (is there a new endpoint that keeps the existing context while adding or streams to the api rather than just from it?)
- tuchsen 3y agoNot associated with this project (or LMQL), but one of the authors of LMQL, a similar project, answered this in a recent thread about it. https://news.ycombinator.com/item?id=35484673#35491123 https://news.ycombinator.com/item?id=35484673#35491123 As a solution to this, we implement speculative execution, allowing us to lazily validate constraints against the generated output, while still failing early if necessary. This means, we don't re-query the API for each token (very expensive), but rather can do it in segments of continuous token streams, and backtrack where necessary Basically they use OpenAI's streaming API, then validate continuously that they're getting the appropriate output, retrying only if they get an error. It's a really clever solution.
- newhouseb 3y agoThis is slick -- It's not explicitly documented anywhere but I hope OpenAI has the necessary callbacks to terminate generation when the API stream is killed rather than continuing in the background until another termination condition happens? I suppose one could check this via looking at API usage when a stream is killed early.
- tuchsen 3y ago
- ahnick 3y agoThis strikes me as being very similar to Jargon (https://github.com/jbrukh/gpt-jargon https://github.com/jbrukh/gpt-jargon), but maybe more formal in its specification?
- ryanklee 3y agoI'm personally starting with learning Guidance and LMQL rather than LangChain just in order to get a better grasp of the behaviors that I've gathered LangChain papers over. Even after that, I'm likely to look at Haystack before LangChain. Just getting the feeling that LangChain is going to end up being considered a kitchen sink solution full of anti patterns so might as well spend time a little lower level while I see which way the winds end up blowing.
- leroy-is-here 3y agoIf this comment performative comedy? Are these real technologies ?
- ryanklee 3y agoNot quite sure what the spirit of your comment is. But, yes, they are real technologies. Very confused as to why you would even find that dubious.
- leroy-is-here 3y agoNot dubious, I just read your comment and it felt like I was reading satire. Even the cadence of your words felt funny. Anyway, I’m not surprised. It’s a new market, everyone’s in on it.
- homarp 3y agoLangChain: https://news.ycombinator.com/item?id=34422627 https://news.ycombinator.com/item?id=34422627 LQML: https://news.ycombinator.com/item?id=35956484 https://news.ycombinator.com/item?id=35956484 Haystack: https://news.ycombinator.com/item?id=29501045 https://news.ycombinator.com/item?id=29501045 or more recently https://news.ycombinator.com/item?id=35430188 https://news.ycombinator.com/item?id=35430188
- WastingMyTime89 3y agoIt is satire. They just don’t realise it yet. It’s pretty clear that we are in the phase where everyone is rushing to get a slice of the pie selling dubious thing and people start parroting word soup hoping they actually make sense and fearing they will miss out. That’s indeed what people often and rightfully satirise about the IT industry. That’s the joke phase before things settle.
- candiddevmike 3y agoWill there be a tool to convert natural language into Guidance?
- lmarcos 3y agoWe can use ChatGPT for that.
- ftxbro 3y agoWill it still be all like "As an AI language model I cannot ..." or can this fix it? I mean asking to sexy roleplay as Yoda isn't the same level as asking how to discreetly manufacture methamphetamine at industrial scale there are levels people
- Der_Einzige 3y agoNo, and in fact I mention that the opposite is the case in the paper I released about constrained text generation: https://paperswithcode.com/paper/most-language-models-can-be-poets-too-an-ai https://paperswithcode.com/paper/most-language-models-can-be... If you ask ChatGPT to generate personal info, say Social Security numbers, it tells you "sorry hal I can't do that". If you constrain it's vocabulary to only allow numbers and hyphens, well, it absolutely will generate things that look like social security numbers, in spite of the instruction tuning. It is for this reason and likely many others that OpenAI does not release the full logits
- alexb_ 3y agoI hope this becomes extremely popular, so that anyone who wants to can completely decouple this from the base model and actually use LLMs to their full potential.
- simonw 3y agoThis is pretty fascinating, but I'm not sure I understand the benefit of using a Handlebars-like DSL here. For example, given this code from https://github.com/microsoft/guidance/blob/main/notebooks/chat.ipynb https://github.com/microsoft/guidance/blob/main/notebooks/ch... create_plan = guidance('''{{#system~}} You are a helpful assistant. {{~/system}} {{#block hidden=True}} {{#user~}} I want to {{goal}}. {{~! generate potential options ~}} Can you please generate one option for how to accomplish this? Please make the option very short, at most one line. {{~/user}} {{#assistant~}} {{gen 'options' n=5 temperature=1.0 max_tokens=500}} {{~/assistant}} {{/block}} {{~! generate pros and cons and select the best option ~}} {{#block hidden=True}} {{#user~}} I want to {{goal}}. ''') How about something like this instead? create_plan = guidance([ system("You are a helpful assistant."), hidden([ user("I want to {{goal}}."), comment("generate potential options"), user([ "Can you please generate one option for how to accomplish this?", "Please make the option very short, at most one line." ]), assistant(gen('options', n=5, temperature=1.0, max_tokens=500)), ]), comment("generate pros and cons and select the best option"), hidden( user("I want to {{goal}}"), ) ])
- jxy 3y agoThey must hate lisp so much that they opt to use {{}} instead.
- armchairhacker 3y agoThe problem with Lisp is that parenthesis are common in regular grammar. {{ is not. Of course input from the user should be escaped, but prompts given by the programmer may have parenthesis and there's no way to disambiguate between the prompt and the DSL.
- evanmoran 3y agoIt's not so much against lisp as double curly is a classic string templating style that is common in web programming. I saw it first with `mustache.js` (first release around 2009), but it's probably been used even before that. https://github.com/janl/mustache.js/ https://github.com/janl/mustache.js/
- Der_Einzige 3y agoThere has been a huge explosion of awesome tooling which utilizes constrained text generation. Awhile ago, I tried my own hand at constraining the output of LLMs. I'm actively working on this to make it better, especially with the lessons learned from repos like this and from guidance https://github.com/hellisotherpeople/constrained-text-generation-studio https://github.com/hellisotherpeople/constrained-text-genera...
- rain1 3y agoThis looks incredible. Wow.
- killthebuddha 3y agoI agree, it looks great. A couple similar projects you might find interesting: - https://github.com/newhouseb/clownfish https://github.com/newhouseb/clownfish - https://github.com/r2d4/rellm https://github.com/r2d4/rellm The first one is JSON only and the second one uses regular expressions, but they both take the same "logit masking" approach as the project GP linked to.
- Der_Einzige 3y agoI love the love from you two - I am trying right now to significantly improve CTGS. I'm not actually using the "Logitsprocessor" from Huggingface, and I really ought to as it will massively speed up inference performance. Unfortunately, fixing up my current code to work with that will take quite awhile. I've started working on it but I am extremely busy these days and would really love for other smart people to help me on this project. If not here, I really want proper access to the constraints APIs (LogitsProcessor and the Constraints classes in Huggingface) in the big webUIs for LLMs like oogabooga. I'd love to make that an extension. I'm also upset at the "undertooling" in the world of LLM prompting. I wrote a snarky blog post about this: https://gist.github.com/Hellisotherpeople/45c619ee22aac6865ca4bb328eb58faf https://gist.github.com/Hellisotherpeople/45c619ee22aac6865c...
- rain1 3y agoDoes this do one query per {{}} thing?
- nico 3y agoIt’s so amazing to see how we are essentially trying to solve “programming human beings” Although on the other hand, that’s what social media and smartphones have already done Maybe AI already took over, doesn’t seem to be wiping out all of humanity
- ubj 3y agoI like this step towards greater rigor when working with LLM's. But part of me can't help but feel like this is essentially reinventing the concept of programming languages: formal and precise syntax to perform specific tasks with guarantees. I wonder where the final balance will end up between the ease and flexibility of everyday language, and the precision / guarantees of a formally specified language.
- intelVISA 3y agoHear me out, just incubated a hot new lang that's about to capture the market and VC hearts: SELECT * FROM llm
- madmax108 3y agoI know you are probably joking, but: https://lmql.ai/ https://lmql.ai/
- lcnPylGDnU4H9OF 3y agoIt won't necessarily turn into some that is fundamentally the same as a current programming language. Rather than a "VM" or "interpreter" or "compiler" we have this "LLM". Even if it requires a lot of domain knowledge to program using an "LLM-interpreted" language, the means of specification (in terms of how the software code is interpreted) may be different enough that it enables easier-to-write, more robust, (more Good Thing) etc. programs.
- davidthewatson 3y agoThis is a hopeful evolutionary path. My concern is that I can literally feel Conway's law emanating from current LLM approaches as they switch between the actual LLM and the governing code around it that layers a buch of conditionals of the form: if (unspeakable_things): return negatory_good_buddy I see this happen a few times per day where the UI triggers a cancel even on its own fake typing mode and overwrites a user response that has at least half-rendered the trigger-warning-inducing response. It's pretty clear from a design perspective that this is intended to be proxy to facial expressions while being worthy of an MVP postmortem discussion about what viability means in a product that's somewhere on a spectrum of unintended consequences that only arise at runtime.
- Animats 3y agoIs this a "language", or just a Python library?
- deleted 3y ago[deleted]
- indus 3y agoThis reminds me of the time when I wrote a cgi script. Basically instructing the templating engine (a very crude regex) to replace session variables, database lookups to the merge fields: Hello {{firstname}}! 1996 and 2023 smells alike.
- hammyhavoc 3y agoRegEx didn't hallucinate though.
- russellbeattie 3y agoThe first 20 versions I write usually do. Make that 50.
- Spivak 3y agoI think it's cool that a company like Microsoft is willing to base a real-boy product on pybars3 which is its author's side-project instead of something like Jinja2. If this catches on I can imagine MS essentially adopting the pybars3 project and turning it into a mature thing.
- mdaniel 3y agoWhich is especially weird given that pybars3 is LGPL and Microsoft prefers MIT stuff
- m3kw9 3y agoI’m not understanding how Guidence Accelerating works. It says “ This cuts this prompt's runtime in half vs. a standard generation approach.” and it gives an example of it asking LLM to generate json. I don’t see anywhere how it accelerates anything because it’s a simple json completion call. How can you accelerate that?
- evanmays 3y agoThe interface makes it look simple, but under the hood it follows a similar approach to jsonformer/clownfish [1] passing control of generation back and forth between a slow LLM and relatively fast python Let's say you're halfway through a generation of a json blob with a name field and a job field and have already generated { "name": "bob" At this point, guidance will take over generation control from the model to generate the next text { "name": "bob", "job": If the model had generated that, you'd be waiting 70 ms per token (informal benchmark on my M2 air). A comma, followed by a newline, followed by "job": is 6 tokens, or 420ms. But since guidance took over, you save all that time. Then guidance passes control back to the model for generating the next field value. { "name": "bob", "job": "programmer" programmer is 2 tokens and the closing " is 1 token, so this took 210ms to generate. Guidance then takes over again to finish the blob { "name": "bob", "job": "programmer" } [1] https://github.com/1rgs/jsonformer https://github.com/1rgs/jsonformer https://github.com/newhouseb/clownfish https://github.com/newhouseb/clownfish Note: guidance is way more general of a tool than these Edit: spacing
- june_twenty 3y agoThanks for that example. Very helpful
- alew1 3y agoBut the model ultimately still has to process the comma, the newline, the "job". Is the main time savings that this can be done in parallel (on a GPU), whereas in typical generation it would be sequential?
- 3y ago
- m3kw9 3y agoThere should be a standard template/language to structurally prompt LLMs. Once that is good, all good LLMs should use the doc to fine tune it to take in that standard. Right now each model has their own little way to best prompt it and you end up needing programs like this to sit in between and handle it for you
- CGamesPlay 3y agoLMQL wants to be that, it seems: https://lmql.ai/ https://lmql.ai/
- bjackman 3y agoWow I think there are details here I'm not fully understanding but this feels like a bit of a quantum leap* in terms of leveraging the strengths while avoiding the weaknesses of LLMs. It seems like anything that provides access to the fuzzy "intelligence" in these systems while minimizing the cost to predictability and efficiency is really valuable. I can't quite put it into words but it seems like we are gonna be moving into a more hybrid model for lots of computing tasks in the next 3 years or so and I wonder if this is a huge peek at the kind of paradigms we'll be seeing? I feel so ignorant in such an exciting way at the moment! That tidbit about the problem solved by "token healing" is fascinating. *I'm sure this isn't as novel to people in the AI space but I haven't seen anything like it before myself.
- Der_Einzige 3y agoA lot of this is because there was and still is systemic undertooling in NLP around how to prompt and leverage the wonderful LLMs that they built. We have to let the Stable Diffusion community guide us, as the waifu generating crowd seems to be quite good at learning how to prompt models. I wrote a snarky github gist about this - https://gist.github.com/Hellisotherpeople/45c619ee22aac6865ca4bb328eb58faf https://gist.github.com/Hellisotherpeople/45c619ee22aac6865c...
- amkkma 3y agoHow does this compare with lmql?
- EddieEngineers 3y agoWhat's with all these weird-looking projects with similar names using Guidance? https://github.com/microsoft/guidance/network/dependents https://github.com/microsoft/guidance/network/dependents They don't even appear to be using Guidance anywhere anyway https://github.com/IFIF3526/aws-memo-server/blob/master/requirements.txt https://github.com/IFIF3526/aws-memo-server/blob/master/requ...
- marcopicentini 3y agoWhat’s the best practice to let an existing Ruby on Rails application use this python framework?
- deleted 3y ago[deleted]
- BeefySwain 3y agoUsing Mustache instead of Jinja for a Python package is a choice
- ianbicking 3y agoI'm having a hard time fully understanding how this works, but I don't think it is simply template substitution. I think it's creating multiple artifacts and completions from the one document. Because of that it's probably much easier if it is a language that can be easily introspected and doesn't support arbitrary expressions.
- BeefySwain 3y agoOkay fair enough then! I'd be interested to see what the rationale was, and if my knee-jerk reaction was unwarranted :)
- imgi456 3y agoDesperate approach from microsoft to gain market share of langchain.
- sheepscreek 3y agoIt is in the same spirit as Maven AI, but takes a slightly different approach. Great to see the progress in this space!
- obiefernandez 3y agoCan this be used with OpenAI APIs?
- forshadowing_ 3y ago[dead]
- iamflimflam1 3y agoVery Interesting. One of the big challenges with LLMs is getting well formed JSON output. GPT4 is much better at this. But is very expensive. So anything that can help is good. Looking forward to trying this out locally with LLAMA.
- kapitanjakc 3y agoWhat is LLAMA ?
- hereforcomments 3y agoFacebook's "leaked" LLM.
- iamflimflam1 3y agoTaks a look at: https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp
- wprl 3y agoThe hubris of mutating spiritual texts as a marketing gimmick is contemptible.
- wahnfrieden 3y agoCould this language be used outside of notebooks?
- jhoffbauer 3y agoHow does this compare to LMQL (https://lmql.ai/ https://lmql.ai/)?