10 ms·
Brex’s Prompt Engineering Guide
- game_the0ry 3y agoI wonder if linguistic and English majors would end up benefiting in this trend of "prompt engineering."
- voz_ 3y agoPerhaps, but this space suffers from the Armageddon astronaut/fire figher problem. It is easier to teach a Computer Science Major good english, than it is to teach an English Major computer science.
- jeron 3y agoI always thought this problem was better thought of as software engineer/farmer problem - it's easer to teach a software engineer about agriculture than the other way around
- xyzzy123 3y agoNow I'm really wondering if that's true or not. Software engineering can be self-studied very cheaply, lots of free resources, mostly you just need time and motivation. Failure is usually cheap. Farming on the other hand requires more local, implicit and hands-on knowledge, capital requirements are high, feedback cycles are slower and failure is expensive.
- dullcrisp 3y agoYeah one doesn’t sound obviously easier than the other to me
- adastra22 3y agoI really, really doubt any of these examples.
- kyleyeats 3y agoThe guys pressing buttons are REALLY confident going into this one.
- voz_ 3y agoI small scale farm (done everything from windowsill to a half acre and back), crops and veggies only, no animals. It’s much more difficult than software engineering, in every way.
- TrueDuality 3y agoI would argue the same is true of farming as well, you're just comparing differences in final outcomes. You can start learning basic farming techniques with a few pots, some seeds, water and a sunny spot in the same way you can get a raspberry pi and a keyboard sourcing information for both from YouTube fairly effectively. There might be an argument for cost at the highest scale in each field which few in either profession really make it to but I'd bet even then its pretty on par. I've been a huge fan of comparing the complexity or cost of a profession. We as a species specialize because doing these tasks in the most efficient way requires managing large number of details that are not obvious at first glance.
- tjr225 3y ago[dead]
- sitkack 3y agoSubstantiate your claim! A farmer is much better generalist than a "software engineer". FFS, the engineer doesn't even use science and the farmer does.
- putnambr 3y agoReally? You think someone who spends all day building systems with software will have an easy time studying and gaining a CDL, learning about soil quality, drainage, fertilization, crop strategies, financing agreements (lease to own agreements, commodities futures, even forex), working with hired labor and managing contractors, engine repair and small scale fabrication? I think the farmer would undoubtedly get bored of being cooped up inside all day at a computer, but they wouldn't have a hard time understanding how to create a system from small moving pieces that they can direct.
- AlotOfReading 3y agoI don't think prompt engineering shares a whole lot with humanities higher education, but the argument that English is easier seems very non-obvious to me. I'd expect the relative skill floor for English to be significantly higher because everyone in the anglosphere gets a minimum of 12 years of intensive English education. Moreover, the starting point for additional higher education is usually native fluency/mastery and even then jobs requiring actual English credentials will usually require a graduate degree on top. Contrast that with programming, where most developers will have learned their first language in college and those 1-3 years of introductory education were enough that they're usually considered hireable by the end of undergrad. Some (few) people can even get to that point with only a 3-6 month intensive bootcamp despite no prior experience.
- sitkack 3y agoCite.
- jatins 3y agoDon't think prompting LLMs require particularly "good" English in the first place. You can say a half baked sentence with typos and it'll still make sense of it. Plus when you go meta and ask LLMs to generate prompts for themselves, you own language proficiency becomes even less important. I do think English/language proficiency will help with Generative image AIs like Midjourney. Like if someone could describe a scene in extreme detail that's more likely to produce a result closer to what you want.
- sharemywin 3y agof*cking Meta. it took me a second to parse your sentence. at first I thought you meant using Llama or something. instead of prompts generating prompts.
- ZephyrBlu 3y agoI actually strongly disagree with this. It's easier to teach a CS major passable English, the bar for which is much lower than passable code. If we cared about English in software engineering as much as we care about code it would be of similar difficulty.
- pram 3y agoWhen Big Data was becoming the hot new thing, I saw people arguing that companies would inevitably need librarians and Masters in Library Science holders etc to wrangle all the information (I guess?) Spoilers: this didn’t happen
- ModernMech 3y agoAbsolutely. My poet wife is able to craft great prompts for Midjourney.
- throwaway675309 3y agoPerhaps this may have been the case early on, but if you observe the trends especially with LLMs moving to zero shot, it's becoming progressively easier to express nuanced instructions with relatively simple prompts.
- owlbite 3y agoFor at least the next few months anyway. The speed this stuff is moving prompt engineering may just be a fad and the next wave of models changes things again. If they persist as a thing, then most of the hard work will just be abstracted away using a standard library of prompts available in a point and click fashion for 99% of use cases.
- bongobingo1 3y agoThats all I think when people espouse the "become a prompt engineer instead!" lines, as if the end goal isn't to remove that exact friction. Otherwise we'd be learning S-LLM-QL instead of "talking" to the bots.
- Veen 3y agoI've occasionally thought that prompt engineering is better described as a type of rhetoric that aims to persuade LLMs to do what you want.
- asteroidz 3y agoThe "Strategies" section looks valuable. Here are a few more great resources from my notes (including one from Lilian Weng who leads Applied Research at OpenAI): - https://lilianweng.github.io/posts/2023-03-15-prompt-engineering https://lilianweng.github.io/posts/2023-03-15-prompt-enginee... - https://www.promptingguide.ai https://www.promptingguide.ai (check the "Techniques" section for several research-vetted approaches) - https://learnprompting.org/docs/intro https://learnprompting.org/docs/intro
- deleted 3y ago[deleted]
- alexbouchard 3y agoYAML is just as effective at communicating data structure to the model while using ~50% less tokens. I now convert all my JSON to YAML before feeding it to GPT API's
- aledalgrande 3y agodo you also get it to return responses in YAML?
- alexbouchard 3y agoYes and then format it back to JSON
- camjw 3y agoI've heard this a lot but don't understand where this idea comes from. With JSON you can strip whitespace whereas with YAML you're stuck with all these pointless whitespace tokens you can't do anything about. I would recommend the exact opposite, JSON is just as effective while using less tokens. This example JSON: {"glossary":{"title":"example glossary","GlossDiv":{"title":"S","GlossList":{"GlossEntry":{"ID":"SGML","SortAs":"SGML","GlossTerm":"Standard Generalized Markup Language","Acronym":"SGML","Abbrev":"ISO 8879:1986","GlossDef":{"para":"A meta-markup language, used to create markup languages such as DocBook.","GlossSeeAlso":["GML","XML"]},"GlossSee":"markup"}}}}} Is 112 tokens, and the corresponding YAML (which I won't paste) is 206. What am I missing?
- thomasfromcdnjs 3y agoI keep going back and fourth between the two. I have absolutely no proof but sometimes feel like the responses I get are weaker if there is no white space in the structured data.
- camjw 3y agoThis is fair, typically I supply data as compact JSON but ask for responses as pretty printed JSON which is quite a large token penalty but tends to strongly reduce malformed JSON outputs.
- ojbyrne 3y ago“In 2017, Google wrote a paper” - there’s the singularity right there.
- uoaei 3y agoThis is a question borne of ignorance: why does Brex, a bank, care about AI like this?
- BoorishBears 3y agoStartups are their bread and butter, Brex spends a lot of money mingling with the space, and AI is currently driving most companies in the space
- uoaei 3y agoAh so it's just marketing? Brex doing a "how do you do, fellow kids" move?
- travisjungroth 3y agoThe release is marketing, both to customers and for hiring. I have zero doubt they made this because they're using LLMs themselves, not just faking it like for fellow kids. LLMs are valuable to all companies of any significant size, just for knowledge management. Then for a bank specifically, there a ton of text classification and summarization tasks when it comes to expense management, bill pay and all the other services they offer. There's also internal stuff. Fraud and KYC would be helped a lot.
- gmuslera 3y agoLets say that you invest in 100 promising technologies. If one of them becomes a black swan, even if all the other 99 failed you still would be winning big.
- morgante 3y agoBrex is launching AI-enabled products: https://www.brex.com/journal/press/brex-openai-ai-tools-for-finance-teams https://www.brex.com/journal/press/brex-openai-ai-tools-for-...
- rvz 3y ago
- typpo 3y agoAre there established best practices for "engineering" prompts systematically, rather than through trial-and-error? Editing prompts is like playing whack-a-mole: once you clear an edge case, a new problem pops up elsewhere. I'd really like to be able to say, "this new prompt performs 20% better across all our test cases". Because I haven't found a better way, I am building https://github.com/typpo/promptfoo https://github.com/typpo/promptfoo, a CLI that outputs a matrix view for quickly comparing outputs across multiple prompts, variables, and models. Good luck to everyone else out there tuning prompts :)
- nico 3y agoAmazing, so useful, thank you
- sitkack 3y agoSeems like you would want to apply some NLP to the prompts themselves Take the gradient of the prompt wrt adjectives, verbs, nouns, etc. I forget the technique, but they add garbage words to the prompt to effectively increase the temperature.
- thomasfromcdnjs 3y agoGreat work, the space needs some more tooling in this direction.
- tlarkworthy 3y agoI use observablehq notebooks so I have programming reactively attached. https://observablehq.com/@tomlarkworthy/colossal-cave-chatgpt-challange https://observablehq.com/@tomlarkworthy/colossal-cave-chatgp...
- ukuina 3y agoThank you so much for this, especially for allowing custom LLM calls to allow testing of local models.
- velavar 3y agoIs it me or is the bot's output in the section "Give a Bot a Fish" incorrect? It states that the most recent receipt is from Mar 5th, 2023 but there are two receipts after that date. This is what worries me about using ChatGPT - the possibility of errors in financial matters, which won't go down well I fear.
- rtsil 3y agoYou're right, and the "Give a bot a fish" method is supposed to be the most reliable. I hope all the apps and platforms that rely on ChatGPT have solid legal disclaimers in place as the liabilities could be quite serious.
- hn_throwaway_99 3y agoThanks very much for posting this! I haven't yet finished reading the whole thing, but even just the first section about the history of LLMs, explaining some of the basic concepts, etc., I found to be a very well-written and useful info, and it was really nice that it linked out to source material. So many times when you go into reading stuff about the latest AI technique or feature it can feel like you need to do a ton of background reading just to understand what they're talking about (especially as the field moves so quickly), so having a nice simple primer at the beginning of this doc was most appreciated!
- nlh 3y ago…was this written by an LLM? I’m starting to doubt anything on the internet that’s overly cherry and polite. Sigh.
- maxbond 3y agoReads like a normal comment to me. Doesn't seem like ChatGPT's typical writing style. It has far too much texture. As someone who people often find stiff and formal, I'm really not looking forward to being accused of not existing more and more often.
- sigstoat 3y ago> I’m starting to doubt anything on the internet that’s overly cherry and polite. > …was this written by an LLM? do you find that asking this is constructive?
- nlh 3y agoI do actually. While I personally am excited about LLMs' potential for good, I think a large swath of the world is ready and equally as excited about their potential for harm / spam / fraud / etc. I'm already seeing bots popping up all over other social channels (Reddit in particular), posting a series of overly-cheery LLM-generated content designed to build up high-karma accounts which then get bought and sold on the not-so-open market so that end users can be further spam'd / defrauded. It sucks. So I'm personally curious in fine-tuning my own "algorithm" for detecting fake content, and pointing it out is helpful for me and, I presume, others who think similarly (and I know there are others). In this case I may absolutely have been wrong, but even the comments above added to my knowledge and helped. So yes, I found it very constructive. I hope others did too.
- jasfi 3y agoI'm working on the idea of features instead of prompts: https://inventai.xyz https://inventai.xyz
- anotherpaulg 3y agoThe suggestion to use markdown tables was quite interesting. It makes a lot of sense, and I haven't seen it described elsewhere. I have been getting good results by asking GPT to produce semi structured responses based on other aspects of (GitHub) markdown. In general, I find it very helpful to find an already popular format that suits your problem. The model is probably already fluent in rendering that output format. So you spend less time trying to teach it the output syntax.
- diarrhea 3y agoIt can do Mermaid diagrams just fine. Very convenient.
- vidarh 3y agoI've even had it generate SVG from Graphviz dot syntax reasonably well, including doing basic layout. It's not great (nowhere near good enough to rely on), but given the complexity of graph layout algorithms that's not surprising. That it can even start to do it and deal with the visuals (e.g. try to avoid overlap etc.) was pretty impressive.
- wearhere 3y agoThis reflects astonishingly poorly on Brex. What customer wants to hear that Brex is using "a non-deterministic model" for "production use cases" like "staying on top of your expenses"? I don't see them acknowledge the downsides of that non-determinism anywhere, let alone hallucination, even though they mention the latter. Hallucinating an extra expense, or missing one, could have serious consequences. This is also potentially terrible from a privacy standpoint. That "staying on top of your expenses" example suggests that you upload "a list of the entire [receipts] inbox" to the model. It _seems_ like they're using OpenAI's API, which doesn’t use customer data for training (unlike ChatGPT), but they should be crystal clear about this. Even if OpenAI doesn't retain/reuse the data, would Brex's customers be happy with this 3rd-party sharing? The expenses example seems like sloppy engineering too—there's no reason to share expense amounts with the model if you just want it to count the number of expenses. Merchant names could be redacted too, replaced with identifiers that Brex would map back to the real data. These suggestions would save on tokens too. Despite Brex saying they're using this in production, I suspect it's mostly a recruiting exercise. It's still a very bad look for their engineering.
- xeyownt 3y ago[flagged]
- vrglvrglvrgl 3y ago[dead]
- zwaps 3y agoWorringly, I am it sure the people working on this really understand what a Transformer is Quote from them: “ There is still active research in non-transformer based language models though, such as Amazon’s AlexaTM 20B which outperforms GPT-3“ Quote from said paper “ For AlexaTM 20B, we used the standard Transformer model architecture“ (Its just an encoder decoder transformer)
- RC_ITR 3y agoYeah, I think the (worrying) confusion is that Amazon calls it a seq2seq model, which was the name of a SOTA RNN from Google a while back. Ofc now, seq2seq just means what you said (an encoder/decoder model, which is actually what a “truly vanilla” transformer would be anyway). The fact that any serious researcher thinks any other serious researchers are using models without self attention is the real red flag here. No one is trying to use other models anymore because they do not scale. There’s enough variety within transformers that you could argue we need a new level of taxonomy, but transformers are basically it for now.
- sgk284 3y agoThanks for pointing this out. That was my mistake – my brain must have swapped out "different transformer architectures" with "different model architectures". I just updated the guide: https://github.com/brexhq/prompt-engineering/commit/3a3ac17a89f3929edae865f2933eade26c63bc28 https://github.com/brexhq/prompt-engineering/commit/3a3ac17a...
- akisej 3y agoThis seems overall well-written and well-explained, but curious for that piece on fine-tuning. This article only recommends it as a last resort. That makes sense for a casual user, but if you're a company seriously using LLMs to provide services for your customers, wouldn't the cost of training data be offset by the potential gains you have and the edge cases you might automatically cover by fine-tuning instead of trying to whack-a-mole predict every single way the prompt can fail?
- ukuina 3y agoThe concern with finetuning, even for specialized use-cases, is that you are binding yourself to the underlying model. Given rapid advancements in the field, this does not seem a prudent use of engineering time. Having a hierarchy of prompts with context stuffing allows for rapid switching across models with a few (non-trivial) surface-level prompt updates while the deeper prompts stay static.
- saladtoes 3y agoI've been playing Gandalf in the last few days, it does a great job at giving an intuition for some of the subtleties of prompt engineering: https://gandalf.lakera.ai https://gandalf.lakera.ai Thanks for putting this together!
- TrueDuality 3y agoWhoa that was a lot of fun. Are you aware of any other games like this? A sort of CTF for AIs?
- mdaniel 3y agohttps://securitycafe.ro/2023/05/15/ai-hacking-games-jailbreak-ctfs/ https://securitycafe.ro/2023/05/15/ai-hacking-games-jailbrea... showed up in my feed this morning but I haven't tried them to know if they're any fun I also found the nondeterministic behavior of Gandalf robbed it of being "fun," to say nothing of the 429s (which they claim to have fixed but I was so burned by the experience I haven't bothered going back through the lower levels to find out)
- dakom 3y agoWhy are we calling this "engineering"? Isn't engineering the application of science to solve problems? (math, definitive logic, etc.) Maybe one day we'll have instruments that let us reason about the connections between prompts and the exact state of the AI, so that we can understand the mechanics of causation, but until then, I would not think that being good at asking questions is "engineering" Are most 10 year olds veteran "search engineers"? Btw I'm asking this slightly tongue-in-cheek, as a discussion point. For example plenty of computer system hacks are done by way of "social engineering", so clearly that term is malleable even within the tech community.
- jollybot 3y agoEngineering, the action of working artfully to bring something about.
- rpastuszak 3y agoWhat's the difference between engineering and design?
- jollybot 3y agoDesign is part of engineering, but engineers will go on to build for example. Many folks like to gatekeep the use of engineer as a word. Mostly it just comes down to the fact that a displine hasn't matured yet and isnt' taught in a rigourous and formal manner. It is still engineering. See Network Engineering for example. Plenty of people building complex systems at scales never before seen. Even more that just about know how traceroute works but still building networks. They are all engineers of a young field of engineering.
- sgk284 3y agoAuthor of the guide here. I attempt to address this in the "Why do we need prompt engineering?"[^1] section. > ... we used an analogy of prompts as the “source code” that a language model “interprets”. Prompt engineering is the art of writing prompts to get the language model to do what we want it to do – just like software engineering is the art of writing source code to get computers to do what we want them to do. Borrowing from Oxford Dictionary, the definition of "engineering" is: > the branch of science and technology concerned with the design, building, and use of engines, machines, and structures. I think it's pretty reasonable to say that "prompt engineering" falls squarely in the realm of "technology concerned with the use of a machine". [^1]: https://github.com/brexhq/prompt-engineering#why-do-we-need-prompt-engineering https://github.com/brexhq/prompt-engineering#why-do-we-need-...
- jaredsohn 3y agoOne thing I haven't heard much discussion about is the fact that ChatGPT is constantly being updated. This means that if you build a prompt for classification and become confident that you've whacked all of the moles so that it is pretty solid with all of the edge cases, it can later start breaking again. Some solutions I can think of are 1) choose a fixed model to test against but they become deprecated over time or 2) perhaps fine-tuning might help.
- tiborsaas 3y agoThat's a valid concern, I think just like with any other software you need to write tests for the AI model to constantly check if your prompts are working as intended. Basic unit tests would work well in this case.
- jaredsohn 3y agoThe differences here compared to unit tests are that the breaking is outside of your control and the updating process is tedious. Also, testing requires making real API calls rather than using stubs so it requires additional infrastructure.
- tiborsaas 3y agoThat's all true, but sometimes things break because some package is bumped. I'd still like to know if my app is basically broken if an LLM has changed somehow. You are right that fine tuning would probably help to minimize the risks, but it probably never can be zero. New tests will also be needed when customers find new edge cases that break our assumptions. Testing LLM prompts is a new paradigm that we'll have to learn to deal with.