17 ms·
Anthropic's Prompt Engineering Tutorial (2024)
- gdudeman 1y agoThis is written for the 3 models (Sonnet, Haiku, Opus 3). While some lessons will be relevant today, others will not be useful or necessary on smarter, RL’d models like Sonnet 4.5. > Note: This tutorial uses our smallest, fastest, and cheapest model, Claude 3 Haiku. Anthropic has two other models, Claude 3 Sonnet and Claude 3 Opus, which are more intelligent than Haiku, with Opus being the most intelligent.
- cjbarber 1y agoYes, Chapters 3 and 6 are likely less relevant now. Any others? Specifically assuming the audience is someone writing a prompt that’ll be re-used repeatedly or needs to be optimized for accuracy.
- babblingfish 1y agoThe big unlock for me reading this is to think about the order of the output. As in, ask it to produce evidence and indicators before answering a question. Obviously I knew LLMs are a probabilistic auto complete. For some reason, I didn't think to use this for priming.
- adastra22 1y agoFurthermore, the opposite behavior is very, very bad. Ask it to give you an answer and justify it, it will output a randomish reply and then enter bullshit mode rationalizing it. Ask it to objectively list pros and cons from a neutral/unbiased perspective and then proclaim an answer, and you’ll get something that is actually thought through.
- beering 1y agoNote that this is not relevant for reasoning models, since they will think about the problem in whatever order it wants to before outputting the answer. Since it can “refer” back to its thinking when outputting the final answer, the output order is less relevant to the correctness. The relative robustness is likely why openai is trying to force reasoning onto everyone.
- adastra22 1y agoThis is misleading if not wrong. A thinking model doesn’t fundamentally work any different from a non-thinking model. It is still next token prediction, with the same position independence, and still suffers from the same context poisoning issues. It’s just that the “thinking” step injects this instruction to take a moment and consider the situation before acting, as a core system behavior. But specialized instructions to weigh alternatives still works better as it ends up thinking about thinking, thinking, then making a choice.
- simianwords 1y agoI think you are misleading as well. Thinking models do recursively generate the final “best” prompt to get the most accurate output. Unless you are genuinely giving new useful information in the prompt, it is kind of useless to structure the prompt in one way or another because reasoning models can generate intermediate steps that give best output. The evidence on this is clear - benchmarks reveal that thinking models are way more performant.
- zurfer 1y agoYou're both kind of right. The order is less important for reasoning models, but if you carefully read thinking traces you'll find that the final answer is sometimes not the same as the last intermediary result. On slightly more challenging problems LLMs flip flop quite a bit and ordering the output cleverly can uplift the result. That might stop being true for newer or future models but I iterated quite a bit in this for sonnet 4.
- stingraycharles 1y agoI typically ask it to start with some short, verbatim quotes of sources it found online (if relevant), as this grounds the context into “real” information, rather than hallucinations. It works fairly well in situations where this is relevant (I recently went through a whole session of setting up Cloudflare Zero Trust for our org, this was very much necessary).
- TingPing 1y agoI try so hard for chatgpt to link and quote real documentation. It makes up links, fake quotes, it even gaslights me when i clarify the information isn’t real.
- axpy906 1y agoSuggest adding 2024 to the title
- 111yoav 1y agoIs there an up to date version of this that was written against their latest models?
- vincnetas 1y agoIt's one year old. Curious how much of it is irrelevant already. Would be nice to see it updated.
- raffael_de 1y agoShould have "(2024)" in the submission title.
- MandieD 1y agoDone
- jwr 1y agoI find the word "engineering" used in this context extremely annoying. There is no "engineering" here. Engineering is about applying knowledge, laws of physics, and rules learned over many years to predictably design and build things. This is throwing stuff at the wall to see if it sticks.
- einrealist 1y agoI call it "Vibe Prompting". Even minor changes to models can render previous prompts useless or invalidate assumptions for new prompts.
- ineedasername 1y agoEven minor changes to a chemical formulation can render previous process design useless or invalidate assumptions for a new formulation. Changing the production or operating process in the face of changing inputs or desired outputs is the bread and butter of countless engineers.
- watwut 1y agoI dont think that is a good argument. In chemical engineering world, the provider who just randomly changes formulations would be called "ureliable, shoody, crappy giving us something we not ordered".
- ineedasername 1y agoIt’s not random. You’re setting up a straw man of what such a change would be. You can arrive at the same chemical formulation through different means. You may also need to tweak a formulation to achieve a variant effect, ie increased resilience of decay from uv light or countless other reasons for tweaks that would necessitate a construction change. It’s the construction I used as the analogous engineering component of this, and that is where an engineer designing the production system would need to understand the changes that the formulation (comparative item to formulation in this metaphor being the model and its underlying components, chemicals, methods of arriving at the end product).
- ipnon 1y agoI really struggle to feel the AGI when I read such things. I understand this is all of year old. And that we have superhuman results in mathematics, basic science, game playing, and other well-defined fields. But why is it difficult to impossible for LLMs to intuit and deeply comprehend what it is we are trying to coax from them?
- xanderlewis 1y ago> superhuman results in mathematics LLMs mostly spew nonsense if you ask them basic questions on research or even master's degree-level mathematics. I've only ever seen non-mathematicians suggest otherwise, and even the biggest mathematician advocate for AI, Terry Tao, seems to recognise this too.
- jimmcslim 1y agoAsk yourself "what is intelligence?". Can intelligence at the level of human experience exist without that which we all also (allegedly) have... "consciousness". What is the source of "consciousness"? Can consciousness be computed? Without answers to these questions, I don't think we are ever achieving AGI. At the end of the day, frontier models are just arithmetic, conditionals, and loops.
- mkl 1y ago> But why is it difficult to impossible for LLMs to intuit and deeply comprehend what it is we are trying to coax from them? It's right there in the name. Large language models model language and predict tokens. They are not trained to deeply comprehend, as we don't really know how to do that.
- kemiller 1y agoHave you ever tried to get an average human to do that? It’s a mixed bag. Computers til now were highly repeatable relative to humans, once programmed, but hopeless at “fuzzy” or associative tasks. Now they have a new trick, that lets them grapple with ambiguity, but the cost is losing that repeatability. The best, most reliable humans were not born that way, it took years or decades of education, and even then it can take a lot of talking to transfer your idea into their brain.
- whatever1 1y agoIn today’s episode of Alchemy for beginners! Reminds me of a time that I found I could speed up by 30% an Algo in a benchmark set if I seed the random number generator with the number 7. Not 8. Not 6. 7.
- ramraj07 1y ago[flagged]
- marcosdumay 1y ago> Like it or not, this IS the job now. Nope. The job is still to come up with working code on the end. If LLMs make your life harder, and you just don't use them, then you'll just get the job done without them.
- whatever1 1y agoI will just enjoy the job security that Seed Science provides me. At any day I could reduce the training costs of all these hyperscalers by 30%. Or maybe not.
- CuriouslyC 1y agoDon't write prompts yourself, use DSPy. That's real prompt "engineering"
- meander_water 1y agoAgree with the other commenters here that this doesn't feel like engineering. However, Anthropic has done some cool work on model interpretability [0]. If that tool was exposed through the public API, then we could at least start to get a feedback loop going where we could compare the internal states of the model with different prompts, and try and tune them systematically. [0] https://www.anthropic.com/research/tracing-thoughts-language-model https://www.anthropic.com/research/tracing-thoughts-language...
- mold_aid 1y ago"Engineering" here seems rhetorically designed to convince people they're not just writing sentences. With respect "prompt writing" probably sounds bad to the same type of person who thinks there are "soft" skills.
- dpe82 1y agoThis strikes me as a silly semantics argument. One could similarly argue software engineering is also just writing sentences with funny characters sprinkled in. Personally, my most productive "software engineering" work is literally writing technical documents (full of sentences!) and talking to people. My mechanical engineering friends report similar as they become more senior.
- watwut 1y agoI dont think so. It says the words were choosen to wngineer peoples emotions and make then feel right way. Tech people do not feel good about "writing propt essay" so it is called engineering to buy their emotional acceptance. Just like we call wrong output "hallucination" rather then "bullshit" or "lie" or "bug" or "wrong output". Hallucination is used to make us feel better and more acceptiong.
- mold_aid 1y ago>This strikes me as a silly semantics argument. Yeah, precisely what I'm saying. I don't think "they write prompt 'engineering' instead of 'writing' to maintain the fragile egos of people who use chatbots" [don't agree? See Mr. "but muh soft skills!" crying down thread] is worth saying outside of HN if I'm honest.
- antinomicus 1y agoThere absolutely are soft skills and it is clear that you do not have them.
- mold_aid 1y agoI mean ok, there's no such thing, so
- Sammi 1y agoHere's my best advice of prompt engineering for hard problems. Always funnel out and then funnel in. Let me explain. State your concrete problem and context. Then we funnel out by asking the AI to do a thorough analysis and investigate all the possible options and approaches for solving the issue. Ask it to go search the web for all possible relevant information. And now we start funneling in again by asking it to list the pros and cons of each approach. Finally we asked it to choose which one or two solutions are the most relevant to our problem at hand. For easy problems you can just skip all of this and just ask directly because it'll know and it'll answer. The issue with harder problems is that if you just ask it directly to come up with a solution then it'll just make something up and it will make up reasons for why it'll work. You need to ground it in reality first. So you do: contrete context and problem, thorough analysis of options, list pros and cons, and pick a winner.
- smallerfish 1y agoNothing about telling it to fuck off, of course to "engineer" its user sentiment analysis?
- taspeotis 1y agoMy workflow has gotten pretty lax around prompts since the models have gotten better. Especially with Claude 4.5 (and 4 before it) once they have a bit of context loaded about the task at hand. I keep it short and conversational, but I do supervise it. If it goes off the rails just smash esc and give it a course correction. And then if you're coming from no context: I throw a bit more detail in at the start and usually start by ending the initial prompt with a question asking it if it can see what I'm talking about in the code; or if it's going to be big: I use planning mode.
- donperignon 1y agoThis AI madness is getting more stupid every day…
- pluc 1y agoSo we've taught this thing how to do what we did and now we need to be taught how to get it to do the things we taught it to do. If this didn't have the entire US economy behind it, it would catch fire like a hot balloon.
- sarcasticsiri 1y agoTheir own example of Claude being able to "reason" is rather funny- https://docs.google.com/spreadsheets/d/1jIxjzUWG-6xBVIa2ay6yDpLyeuOh_hR_ZB75a47KX_E/edit?gid=2055375080#gid=2055375080&range=E40:E43 https://docs.google.com/spreadsheets/d/1jIxjzUWG-6xBVIa2ay6y... First it says that if you ask a logic problem without any system prompt it will not get it. But if you add a system prompt that "You are a logic bot designed to answer complex logic problems.", it will get it right. It gets "a married person (Jack) is looking at an unmarried person (George)" which is simply incorrect! They also add "Although notably not for all the right reasons", is it to justify the logic?
- deleted 1y ago[deleted]
- kuharich 1y agoPast comments: https://news.ycombinator.com/item?id=41395921 https://news.ycombinator.com/item?id=41395921
- kgeist 1y agoYesterday I was trying to make a small quantized model work, but it just refused to follow all my instructions. I tried to use all the tricks I could remember, but fixing instruction-following for one rule would always break another. Then I had an idea: do I really want to be a "prompt engineer" and waste time on this, when the latest SOTA models probably already have knowledge of how to make good prompts in their training data? Five minutes and a few back-and-forths with GPT-5 later, I had a working prompt that made the model follow all my instructions. I did it manually, but I'm sure you can automate this "prompt calibration" with two LLMs: a prompt rewriter and a judge in a loop.
- pojzon 1y agoThats how copilot works by default. At least in IDE, it takes my prompt, makes it pretty and passes it further.
- b0dhimind 1y agoAre there any project-based up-to-date guides to Agentic coding w/ VS Code? Including all the latest features they keep adding from Cursor and other IDEs. Since we can just bring our own free models there.