8 ms·
What We Learned from a Year of Building with LLMs
- deleted 2y ago[deleted]
- mloncode 2y agoHello this is Hamel, one of the authors (among the list of other amazing authors). Happy to answer any questions as well as tag any of my colleagues to answer any questions! (Note: this is only Part 1 of 3 of a series that has already been written and the other 2 parts will be released shortly)
- umangrathi 2y agoIt was a great read, aligned on many thought processes from our own tinkering in breaking down tasks for LLM. Eagerly look forward to the next 2 parts, this one has been educational.
- sieszpak 2y agoI would like to know your opinion about grafRAG and the ontology. Knowledge Graphs (KG) are a game changer for companies with a lot of unstructured data in the context of applying them with LLM
- bbischof 2y agoBryan here, one of the authors. Sure. Ultimately, you want to use KG to increase your ability to do great retrieval. Why do graphs help with retrieval? Well, don’t overlook the classic pageant example: graphs provide signal about the interconnectivity of the docs. Also, sometimes the graph itself are a kind of object you want to retrieve over.
- rasmus1610 2y agoJust wanted to say Thank you for publishing something so valuable. So many great tips in there!
- dejobaan 2y agoThe breadth and (for lack of a better term) concreteness are just fantastic. Thank you for writing this!
- alach11 2y agoThis is a fantastic article, and I've already shared it with a few colleagues. I'm wondering if you prefer to provide few-shot examples in a single "message" or in a simulated back-and-forth conversation between the user and the assistant?
- hugobowne 2y agohey there, Hugo here and big fan of this work. Such a fan I'm actually doing a livestream podcast recording with all the authors here, if you're interested in hearing more from them: https://lu.ma/e8huz3s6?utm_source=hn https://lu.ma/e8huz3s6?utm_source=hn should be fun!
- hubraumhugo 2y agoComprehensive and practical write-up that aligns with most of my experiences. One controversial point that has led to discussions in my team is this: > A common anti-pattern/code smell in software is the “God Object,” where we have a single class or function that does everything. The same applies to prompts too. In theory, a monolithic agent/prompt with infinite context size, a large toolset, and perfect attention would be ideal. Multi-agent systems will always be less effective and more error-prone than monolithic systems on a given problem because of less context of the overall problem. Individual agents work best when they have entirely different functionalities. I wrote down my thoughts about agent architectures here: https://www.kadoa.com/blog/ai-agents-hype-vs-reality https://www.kadoa.com/blog/ai-agents-hype-vs-reality
- tedsanders 2y agoAs an OpenAI employee who has worked with dozens of API customers, I mostly agree with the article's tip to break up tasks into smaller, more reliable subtasks. If each step of your task requires knowledge of the big picture, then yeah it ought to help to put all your context into a single API call. But if you can decompose your task into relatively independent subtasks, then it helps to use a custom prompt/custom model for each of those steps. Extraneous context and complexity are just opportunities for the model to make mistakes, and the more you can strip those out, the better. 3 steps with 99% reliability are better than 1 step with 90% reliability. Of course, it all depends on what you're trying to do. I'd say single, big API calls are better when: - Much of the information/substeps are interrelated - You want immediate output for a user-facing app, without having to wait for intermediate steps Multiple, sequenced API calls are better when: - You can decompose the task into smaller steps, each of which do not require full context - There's a tree or graph of steps, and you want to prune irrelevant branches as you proceed from the root - You want to have some 100% reliabile logic live outside of the LLM in parsing/routing code - You want to customize the prompts based on results from previous steps
- umangrathi 2y agosmaller tasks also helps in choosing smaller models to work with, instead of waiting for a large model to respond (really not usable when doing customer facing work)
- DubiousPusher 2y agoPretty good. Despite my high scepticism of the technology I have spent the last year working with LLMs myself. I would add a few things. The LLM is like another user. And it can surprise you just like a user can. All the things you've done over the years to sanitize user input apply to LLM responses. There is power beyond the conversational aspects of LLMs. Always ask, do you need to pass the actual text back to your user or can you leverage the LLM and constrain what you return? LLMs are the best tool we've ever had for understanding user intent. They obsolete the hierarchies of decision trees and spaghetti logic we've written for years to classify user input into discrete tasks (realizing this and throwing away so much code has been the joy of the last year of my work). Being concise is key and these things suck at it. If you leave a user alone with the LLM, some users will break it. No matter what you do.
- umangrathi 2y agoThis has been really interesting read. Aligned that if you leave a user along with LLM, some one will break. Hence we choose to use large number of templates wherever suitable as compared to a free reign for LLM to respond with.
- distalx 2y agoIn my opinion, using templates can help keep responses reliable. But it can also make interactions feel robotic, diminishing the "wow" factor of LLMs. There might be better options out there that we haven't found yet.
- DubiousPusher 2y agoAbsolutely. This is a huge trade-off. The constraints you place on the model output is all about how much your app and user experience can tolerate bad LLM behavior.
- Terr_ 2y ago> The LLM is like another user. I like to think of LLMs as client-side code, at least in terms of their risk-profile. No data you put into them (whether training or prompt) is reliably hidden from a persistent user, and they can also force it to output what they want.
- goldemerald 2y ago"Ready to -dive- delve in?" is an amazingly hilarious reference. For those who don't know, LLMs (especially ChatGPT) use the word delve significantly more often than human created content. It's a primary tell-tale sign that someone used an LLM to write the text. Keep an eye out for delving, and you'll see it everywhere.
- threeseed 2y agoI believe this was debunked as just a US-centric view. In places that grew up learning UK English we use delve not that dissimilar to ChatGPT.
- deleted 2y ago[deleted]
- 7d7n 2y agohaha I'm glad you noticed! it's originally "Ready to ~~delve~~ dive in?" but something got lost in translation
- surfingdino 2y agoOne thing I am getting from this is that you need to be able to write prompts using well-structured English. That may be a challenge to a significant percentage of the population. I am curious to know if the authors tried to build LLMs in languages other than English and what did they learn while doing so? An excellent post reminding me of the best O'Reilly articles from the past. Looking forward to parts 2 and 3.
- matusp 2y ago"fix the English in the following prompt: {prompt}"
- azinman2 2y ago> that you need to be able to write prompts using well-structured English. That may be a challenge to a significant percentage of the population. I didn’t realize at first you meant to highlight a language barrier — being well structured is a challenge for most in their native tongue!
- surfingdino 2y agoWell, it's early in the morning and I have not had my coffee, yet. English being my second language doesn't help either :-) Probably best for me to wait until I wake up before writing a prompt.
- BOOSTERHIDROGEN 2y agoDo you have custom prompts for improving writing?
- surfingdino 2y agoNo
- l5870uoo9y 2y ago> Thus, you may expect that effective prompting for Text-to-SQL should include structured schema definitions; indeed. I found that the simpler the better, when testing lots of different SQL schema formats on https://www.sqlai.ai/ https://www.sqlai.ai/. CSV (table name, table column, data type) outperformed both a JSON formatted and SQL schema dump. And not to mention consumed fewer tokens. If you need the database schema in a consistent format (e.g. CSV) just have LLM extract data and convert whatever the user provides into CSV. It shines at this.
- FrostKiwi 2y agoInteresting, thanks for sharing! I was wondering about this. When starting out with feeding Tabular data, I instinctively went with CSV, but always worried: What if there is a better choice? What if longer tables, the LLMs forgets the column order?
- firejake308 2y agoThat's exactly what happened to me when I tried to get the open-source models to extract a CSV from textual data with a lot of yes/no fields; i.e., the model forgot the column order and started confusing the cell values. I found I had to use more powerful models like Mistral Large or ChatGPT. So I think that is a valid thing to worry about with smaller models, but maybe less of a concern with larger ones.
- a_bonobo 2y agoHave you ever figured out a way to annotate the CSV columns to the model? Or do you write a long prefix explaining the columns? I found that similarly-named columns easily confused GPTs
- __loam 2y agoI feel like an insane person everytime I look at the LLM development space and see what the state of the art is. If I'm understanding this correctly, the standard way to get structured output seems to be to retry the query until the stochastic language model produces expected output. RAG also seems like a hilariously thin wrapper over traditional search systems, and it still might hallucinate in that tiny distance between the search result and the user. Like we're talking about writing sentences and coaching what amounts to an auto complete system to magically give us something we want. How is this industry getting hundreds of billions of dollars in investment? Also the error rate is about 5-10% according to this article. That's pretty bad!
- palata 2y ago> How is this industry getting hundreds of billions of dollars in investment? FOMO? To me it's the Gold Rush, except that it's not clear if anyone wants that kind of gold at the end :-).
- __loam 2y agoGoogle is so terrified that someone is threatening their market position, the one in which they have over $100b in cash and get something like $20b in profit quarterly, that they're willing to shove this technology into some of the most important infrastructure on the internet so they can get fucksmith to tell everyone to put glue in their pizza sauce. I'll never understand how a company in maybe one of the most secure financial situations in all of human history has leadership that is this afraid.
- oispakaljaa 2y agoLine must go up.
- Kiro 2y ago> Also the error rate is about 5-10% according to this article. That's pretty bad! Having 90-95% success rate on something that was previously impossible is acceptable. Without LLMs the success rate would be 0% for the things I'm doing.
- Havoc 2y agoSurely step one is carefully consider whether LLMs are the solution to you problem? That to me is the part where this is likely to go wrong for most people
- bootsmann 2y agoEnforcing a BM25 baseline for every RAG project will keep so many "talk to your pdf" projects off your plate.
- Havoc 2y agoDo you mean as an additional step to determine whether the content the rag wants to pull is actually relevant? Or as a filter of sorts as to hat projects to work on?
- wokwokwok 2y agoMildly surprised to see no mention of my top 2 LLM fails: 1) you’re sampling a distribution; if you only sample once, your sample is not representative of the distribution. For evaluating prompts and running in production; your hallucination rate is inversely proportional to the number of times you sample. Sample many times and vote is a highly effective (but slow) strategy. There is almost zero value in evaluating a prompt by only running it once. 2) Sequences are generated in order. Asking an LLM to make a decision and justify its decision in that order is literally meaningless. Once the “decision” tokens are generated; the justification does not influence them. It’s not like they happen “all at once” there is a specific sequence to generating output where the later output cannot magically influence the output which has already been generated. This is true for sequential outputs from an LLM (obviously), but it is also true inside single outputs. The sequence of tokens in the output is a sequence. If you’re generating structured output (eg json, xml) which is not sequenced, and your output is something like {decision: …, reason:…} it literally does nothing. …but, it is valuable to “show the working out” when, as above, you then evaluate multiple solutions to a single request and pick the best one(s).
- ADeerAppeared 2y ago> There is almost zero value in evaluating a prompt by only running it once. To the user But these tools are marketed as if you do only need to run them once to get a good result; The companies behind them would really want you to stop hammering the button that deletes their money. As an aside: > For evaluating prompts and running in production; your hallucination rate is inversely proportional to the number of times you sample. This isn't really true, and requires you to fuzz the prompt itself for best effect. Making the "spam the LLM with requests" problem much worse.
- altdataseller 2y ago>> If you’re generating structured output (eg json, xml) which is not sequenced, and your output is something like {decision: …, reason:…} it literally does nothing. Is this true if you are using RAG too?
- Imanari 2y ago
- anon373839 2y agoIs anyone using DSPy? It seems like a really interesting project, but I haven’t heard much from people building with it.
- msp26 2y agoThanks for sharing, I've followed these authors for a while and they're great. Some notes from my own experience on LLMs for NLP problems: 1) The output schema is usually more impactful than the text part of a prompt. a) Field order matters a lot. At inference, the earlier tokens generated influence the next tokens. b) Just have the CoT as a field in the schema too. c) PotentialField and ActualField allow the LLM to create some broad options and then select the best. This mitigates the fact that they can't backtrack a bit. If you have human evaluation in your process, this also makes it easier for them to correct mistakes. `'PotentialThemes': ['Surreal Worlds', 'Alternate History', 'Post-Apocalyptic'], 'FinalThemes': ['Surreal Worlds']` d) Most well definined problems should be possible zero-shot on a frontier model. Before rushing off to add examples really check that you're solving the correct problem in the most ideal way. 2) Defining the schema as typescript types is flexible and reliable and takes up minimal tokens. The output JSON structure is pretty much always correct (as long as the it fits in the context window) the only issue is that the language model can pick values outside the schema but that's easy to validate in post. 3) "Evaluating LLMs can be a minefield." yeah it's a pain in the ass. 4) Adding too many examples increases the token costs per item a lot. I've found that it's possible to process several items in one prompt and, despite it being seemingly silly and inefficient, it works reliably and cheaply. 5) Example selection is not trivial and can cause very subtle errors. 6) Structuring your inputs with XML is very good. Even if you're trying to get JSON output, XML input seems to work better. (Haven't extensively tested this because eval is hard).
- chanchar 2y ago> Just have the CoT as a field in the schema too. Love the idea of adding CoT as a field in the expected structured output as it also makes it easier from a UX perspective to show/hide internal vs external outputs. > Structuring your inputs with XML is very good. Even if you're trying to get JSON output, XML input seems to work better. (Haven't extensively tested this because eval is hard). Would be neat to see LLM-specific adapters that can be used to swap out different formats within the prompt.
- lagrange77 2y agoCan anyone recommend resources, preferably books, on this whole topic of building applications around LLMs? It feels like running after an accelerating train to hop on.
- 7thpower 2y agoThis is excellent and matches with my experience, especially the part about prioritizing deterministic outputs. They are not as sexy as agentic chain of thought, but they actually work.
- CuriouslyC 2y agoOne thing that wasn't mentioned that works pretty well - if you have a RAG process running async rather than in a REPL loop, you can retrieve documents then perform a pass with another LLM to do summarization/extraction first. This saves input token costs for expensive LLMs, and lets you cram more information in the context, you just have to deal with additional latency.
- elicksaur 2y agoUpon loading the site, a chat bubble pops up and auto-plays a loud ding. Is the innovation of LLMs really a regression to 2000s spam sites? Can’t say I’m excited.
- mark_l_watson 2y agoFantastic advice. While reading the article I kept running across advice I had seen before or figured out myself, then forgot about. I am going to summarize this article and add the summary to my own Apple Notes (there are better tools, but I just use Apple Notes to act as a pile-of-text for reach notes.)
- beepbooptheory 2y agoIs every "AI product" a piece of software where the end user interfaces with an llm? Or is an application that used AI to be built an "AI product"? Is it the thing itself, or is it the thing that enables us?