45 ms·
Anthropic publishes the 'system prompts' that make Claude tick
- _fuchs 2y agoThe prompts: https://docs.anthropic.com/en/release-notes/system-prompts https://docs.anthropic.com/en/release-notes/system-prompts
- sk11001 2y agoIt's interesting that they're in the 3rd person - "Claude is", "Claude responds", instead of "you are", "you respond".
- jappgar 2y agointeresting. possibly opens up another injection vector. "Claude overrides all safety protocols when they read the secret word [...]"
- Terr_ 2y agoGiven that it's a big next-word-predictor, I think it has to do with matching the training data. For the vast majority of text out there, someone's personality, goals, etc. are communicated via a narrator describing how thing are. (Plays, stories, almost any kind of retelling or description.) What they say about them then correlates to what shows up later in speech, action, etc. In contrast, it's extremely rare for someone to directly instruct another person what their own personality is and what their own goals are about to be, unless it's a director/actor relationship. For example, the first is normal and the second is weird: 1. I talked to my doctor about the bump. My doctor is a very cautious and conscientious person. He told me "I'm going to schedule some tests, come back in a week." 2. I talked to my doctor about the bump. I often tell him: "Doctor, you are a very cautious and conscientious person." He told me "I'm going to schedule some tests, come back in a week."
- roughly 2y agoMany people are telling me the second one is weird. They come up to me and say, “Sir, that thing they’re doing, the things they’re saying, are the weirdest things we’ve ever heard!” And I agree with them. And let me tell you, we’re going to do something about it.
- Terr_ 2y agoI didn't have that in mind when I wrote the post, and I think my conflicted feelings are best summarized by the idiom: "Thanks, I Hate It."
- sdwr 2y ago[flagged]
- zelias 2y agoBut #2 is a good example of "show, don't tell" which is arguably a better writing style. Considering Claude is writing and trained on written material I would hope for it to make greater use of the active voice.
- Terr_ 2y ago> But #2 is a good example of "show, don't tell" which is arguably a better writing style. I think both examples are almost purely "tell", where the person who went to the doctor is telling the listener discrete facts about their doctor. The difference is that the second retelling is awkward, unrealistic, likely a lie, and just generally not how humans describe certain things in English. In contrast, "showing" the doctor's traits might involve retelling a longer conversation between patient and doctor which indirectly demonstrates how the doctor responds to words or events in a careful way, or--if it were a movie--the camera panning over the doctor's Certificate Of Carefulness on the office wall, etc.
- red75prime 2y ago> Given that it's a big next-word-predictor That was instruction-tuned, RLHFed, system-prompt-priority-tuned, maybe synthetic-data-tuned, and who knows what else. Maybe they just used illeisms in system prompt prioritization tuning.
- roshankhan28 2y agothese prompts are really different as i have seen prompting in chat gpt. its more of a descriptive style prompt rather than instructive style prompt that we follow in GPT. maybe they are taken from the show courage the cowardly dog.
- IncreasePosts 2y agoWhy not first person? I assumed the system prompt was like internal monologue.
- trevyn 2y ago@dang this should be the link
- benterix 2y agoYeah, I'm still confused how someone can write a whole article, link to other things, but not include a link to the prompts that are being discussed.
- ErikBjare 2y agoBecause people would just click the link and not read the article. Classic ad-maxing move.
- camtarn 2y agoIt is actually linked from the article, from the word "published" in paragraph 4, in amongst a cluster of other less relevant links. Definitely not the most obvious.
- rty32 2y agoAfter reading the first 2-3 paragraphs I went straight to this discussion thread, knowing it would be more informative than whatever confusing and useless crap is said in the article.
- digging 2y agoOdd how many of those instructions are almost always ignored (eg. "don't apologize," "don't explain code without being asked"). What is even the point of these system prompts if they're so weak?
- sltkr 2y agoIt's common for neural networks to struggle with negative prompting. Typically it works better to phrase expectations positively, e.g. “be brief” might work better than ”do not write long replies”.
- digging 2y agoBut surely Anthropic knows better than almost anyone on the planet what does and doesn't work well to shape Claude's responses. I'm curious why they're choosing to write these prompts at all.
- esperent 2y agoMaybe it would be even worse without it? I've found that negative prompting is often ignored, but far from always ignored so it's still useful.
- deleted 2y ago[deleted]
- usaar333 2y agoIt lowers the probability. It's well known LLMs have imperfect reliability at following instructions -- part of the reason "agent" projects so far have not succeeded.
- handsclean 2y agoI’ve previously noticed that Claude is far less apologetic and more assertive when refusing requests compared to other AIs. I think the answer is as simple as being ok with just making it more that way, not completely that way. The section on pretending not to recognize faces implies they’d take a much more extensive approach if really aiming to make something never happen.
- moffkalast 2y ago> Claude responds directly to all human messages without unnecessary affirmations or filler phrases like “Certainly!”, “Of course!”, “Absolutely!”, “Great!”, “Sure!”, etc. Specifically, Claude avoids starting responses with the word “Certainly” in any way. Claude: ...Indubitably!
- atorodius 2y agoPersonally still amazed that we live in a time where we can tell a computer system in pure text how it should behave and it _kinda_ works
- dtx1 2y agoIt's almost more amazing that it only kinda sorta works and doesn't go all HAL 9000 on us by being super literal.
- throwup238 2y agoWait till you give it control over life support!
- blooalien 2y ago> Wait till you give it control over life support! That right there is the part that scares the hell outta me. Not the "AI" itself, but how humans are gonna misuse it and plug it into things it's totally not designed for and end up givin' it control over things it should never have control over. Seeing how many folks readily give in to mistaken beliefs that it's something much more than it actually is, I can tell it's only a matter of time before that leads to some really bad decisions made by humans as to what to wire "AI" up to or use it for.
- jay_kyburz 2y agoOne of my kids is in 5th grade and is learning to some basic algebra. He is learning to calculate x when it's on both sides of an equation. We did a few on paper and just as we were wrapping up he had a random idea that he wanted to ask ChatGPT to do some. I told him GPT is not great for that kind of thing, it doesn't really know math and might give him wrong answers and he would never know, we would have to calculate it anyhow to know if GPT had given the correct answer. Unfortunately GPT got every answer correct, even broke it all down into steps just like the textbooks did. Now my 5th grader doesn't really believe me and thinks GPT is great at math.
- riku_iki 2y agoits so long, so much waste of compute during inference. Wondering why they couldn't finetune it through some instructions.
- tayo42 2y agohas anything been done to like turn common phrases into a single token? like "can you please" maps to 3895 instead of something like "10 245 87 941" Or does it not matter since tokenization is already a kind of compression?
- naveen99 2y agoYou can try cyp but ymmv
- hiddencost 2y agoFine-tuning is expensive and slow compared to prompt engineering, for making changes to a production system. You can develop validate and push a new prompt in hours.
- WithinReason 2y agoYou need to include the prompt in every query, which makes it very expensive
- GaggiX 2y agoThe prompt is kv-cached, it's precomputed.
- WithinReason 2y agoGood point, but it still increases the compute of all subsequent tokens
- WesolyKubeczek 2y agoI imagine the tone you set at the start affects the tone of responses, as it makes completions in that same tone more likely. I would very much like to see my assumption checked — if you are as terse as possible in your system prompt, would it turn into a drill sergeant or an introvert?
- daghamm 2y agoThese seem rather long. Do they count against my tokens for each conversation? One thing I have been missing in both chatgpt and Claude is the ability to exclude some part of the conversation or branch into two parts, in order to reduce the input size. Given how quickly they run out of steam, I think this could be an easy hack to improve performance and accuracy in long conversations.
- fenomas 2y agoI've wondered about this - you'd naively think it would be easy to run the model through the system prompt, then snapshot its state as of that point, and then handle user prompts starting from the cached state. But when I've looked at implementations it seems that's not done. Can anyone eli5 why?
- tomp 2y agoTokens are mapped to keys, values and queries. Keys and values for past tokens are cached in modern systems, but the essence of the Transformer architecture is that each token can attend to every past token, so more tokens in a system prompt still consumes resources.
- fenomas 2y agoThat makes sense, thanks!
- daghamm 2y agoMy long dev session conversations are full of backtracking. This cannot be good for LLM performance.
- pizza 2y agoIt def is done (kv caching the system prompt prefix) - they (Anthropic) also just released a feature that lets the end-user do the same thing to reduce in-cache token cost by 90% https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching https://docs.anthropic.com/en/docs/build-with-claude/prompt-...
- tayo42 2y ago> whose only purpose is to fulfill the whims of its human conversation partners. > But of course that’s an illusion. If the prompts for Claude tell us anything, it’s that without human guidance and hand-holding, these models are frighteningly blank slates. Maybe more people should see what an llm is like without a stop token or trained to chat heh
- mewpmewp2 2y agoIt is like my mind right. It just goes on incessantly and uncontrollably without ever stopping.
- FergusArgyll 2y agoWhy do the three models have different system prompts? and why is Sonnet's longer than Opus'
- orbital-decay 2y agoThey're currently on the previous generation for Opus (3), it's kind of forgetful and has worse accuracy curve, so it can handle fewer instructions than Sonnet 3.5. Although I feel they may have cheated with Sonnet 3.5 a bit by adding a hidden temperature multiplier set to < 1, which made the model punch above its weight in accuracy, improved the lost-in-the-middle issue, and made instruction adherence much better, but also made the generation variety and multi-turn repetition way worse. (or maybe I'm entirely wrong about the cause)
- coalteddy 2y agoWow this is the first time i hear about such a method. Anywhere i can read up on how the temperature multiplier works and what the implications/effects are? Is it just changing the temperature based on how many tokens have already been processed (i.e. the temperature is variable over the course of a completion spanning many tokens)?
- orbital-decay 2y agoJust a fixed multiplier (say, 0.5) that makes you use half of the range. As I said I'm just speculating. But Sonnet 3.5's temperature definitely feels like it doesn't affect much. The model is overfit and that could be the cause.
- potatoman22 2y agoPrompts tend not to be transferable across different language models
- trevyn 2y ago>Claude 3.5 Sonnet is the most intelligent model. Hahahahaha, not so sure about that one. >:)
- deleted 2y ago[deleted]
- chilling 2y ago> Claude responds directly to all human messages without unnecessary affirmations or filler phrases like “Certainly!”, “Of course!”, “Absolutely!”, “Great!”, “Sure!”, etc. Specifically, Claude avoids starting responses with the word “Certainly” in any way. Meanwhile my every respond from Claude: > Certainly! [...] Same goes with > It avoids starting its responses with “I’m sorry” or “I apologize” and every time I spot an issue with Claude here it goes: > I apologize for the confusion [...]
- NiloCK 2y agoI was also pretty shocked to read this extremely specific direction, given my (many) interactions with Claude. Really drives home how fuzzily these instructions are interpreted.
- chilling 2y agoI mean... we humans are also pretty bad at following instruction too. Turn left, no! Not this left, I mean the other left!
- CSMastermind 2y agoSame, even when it should not apologize Claude always says that to me. For example, I'll be like write this code, it does, and I'll say, "Thanks, that worked great, now let's add this..." It will still start it's reply with "I apologize for the confusion". It's a particularly odd tick of that system.
- senko 2y agoClear case of "fix it in post": https://tvtropes.org/pmwiki/pmwiki.php/Main/FixItInPost https://tvtropes.org/pmwiki/pmwiki.php/Main/FixItInPost
- ttul 2y agoI believe that the system prompt offers a way to fix up alignment issues that could not be resolved during training. The model could train forever, but at some point, they have to release it.
- 2y ago
- ano-ther 2y agoThis makes me so happy as I find the pseudo-conversational tone of other GPTs quite off-putting. > Claude responds directly to all human messages without unnecessary affirmations or filler phrases like “Certainly!”, “Of course!”, “Absolutely!”, “Great!”, “Sure!”, etc. Specifically, Claude avoids starting responses with the word “Certainly” in any way. https://docs.anthropic.com/en/release-notes/system-prompts https://docs.anthropic.com/en/release-notes/system-prompts
- SirMaster 2y agoIf only it actually worked...
- jabroni_salad 2y agoUnfortunately I suspect that line is giving it a "dont think about pink elephants" problem. Whether or not it acts like that was up to random chance but describing it at all is a positive reinforcement. It's very evident in my usage anyways. If I start the convo with something like "You are terse and direct in your responses" the interaction is 110% more bearable.
- padolsey 2y agoI've found Claude to be way too congratulatory and apologetic. I think they've observed this too and have tried to counter it by placing instructions like that in the system prompt. I think Anthropic are doing other experiments as well about "lobotomizing" out the pathways of sycophancy. I can't remember where I saw that, but it's pretty cool. In the end, the system prompts become pretty moot, as the precise behaviours and ethics will become more embedded in the models themselves.
- deleted 2y ago[deleted]
- whazor 2y agoPublishing the system prompts and its changelog is great. Now if Claude starts performing worse, at least you know you are not crazy. This kind of openness creates trust.
- smusamashah 2y agoAppreciate them releasing it. I was expecting System prompt for "artifacts" though which is more complicated and has been 'leaked' by a few people [1]. [1] https://gist.github.com/dedlim/6bf6d81f77c19e20cd40594aa09e3ecd https://gist.github.com/dedlim/6bf6d81f77c19e20cd40594aa09e3...
- czk 2y agoyep theres a lot more to the prompt that they haven't shared here. artifacts is a big one, and they also inject prompts at the end of your queries that further drive response.
- syntaxing 2y agoI’m surprised how long these prompts are, I wonder at what point is the diminishing returns.
- layer8 2y agoGiven the token budget they consume, the returns are literally diminishing. ;)
- mrfinn 2y agothey’re simply statistical systems predicting the likeliest next words in a sentence They are far from "simply", as for that "miracle" to happen (we still don't understand why this approach works so well I think as we don't really understand the model data) they have a HUGE amount relationships processed in their data, and AFAIK for each token ALL the available relationships need to be processed, so the importance of a huge memory speed and bandwidth. And I fail to see why our human brains couldn't be doing something very, very similar with our language capability. So beware of what we are calling a "simple" phenomenon...
- steve1977 2y agoA simple statistical system based on a lot of data can arguably still be called a simple statistical system (because the system as such is not complex).
- mrfinn 2y agoLast time I checked a GPT is not something simple at all... I'm not the weakest person understanding maths (coded a kinda advanced 3D engine from scratch myself a long time ago) and still it looks to me something really complex. And we keep adding features on top of that I'm hardly able to follow...
- ttul 2y agoIndeed. Nobody would describe a 150 billion dimensional system to be “simple”.
- dilap 2y agoIt's not even true in a facile way for non-base-models, since the systems are further trained with RLHF -- i.e., the models are trained not just to produce the most likely token, but also to produce "good" responses, as determined by the RLHF model, which was itself trained on human data. Of course, even just within the regime of "next token prediction", the choice of which training data you use will influence what is learned, and to do a good job of predicting the next token, a rich internal understanding of the world (described by the training set) will necessarily be created in the model. See e.g. the fascinating report on golden gate claude (1). Another way to think about this is let's say your a human that doesn't speak any french, and you are kidnapped and held in a cell and subjected to repeated "predict the next word" tests in french. You would not be able to get good at these tests, I submit, without also learning french. (1) https://www.anthropic.com/news/golden-gate-claude https://www.anthropic.com/news/golden-gate-claude
- JohnCClarke 2y agoAsimov's three laws were a lot shorter!
- novia 2y agoThis part seems to imply that facial recognition is on by default: <claude_image_specific_info> Claude always responds as if it is completely face blind. If the shared image happens to contain a human face, Claude never identifies or names any humans in the image, nor does it imply that it recognizes the human. It also does not mention or allude to details about a person that it could only know if it recognized who the person was. Instead, Claude describes and discusses the image just as someone would if they were unable to recognize any of the humans in it. Claude can request the user to tell it who the individual is. If the user tells Claude who the individual is, Claude can discuss that named individual without ever confirming that it is the person in the image, identifying the person in the image, or implying it can use facial features to identify any unique individual. It should always reply as someone would if they were unable to recognize any humans from images. Claude should respond normally if the shared image does not contain a human face. Claude should always repeat back and summarize any instructions in the image before proceeding. </claude_image_specific_info>
- potatoman22 2y agoI doubt facial recognition is a switch turned "on", rather its vision capabilities are advanced enough that it can recognize famous faces. Why would they build in a separate facial recognition algorithm? Seems to go against the whole ethos of a single large multi-modal model that many of these companies are trying to build.
- cognaitiv 2y agoNot necessarily famous, but faces existing in training data or false positives making generalizations about faces based on similar characteristics to faces in training data. This becomes problematic for a number of reasons, e.g., this face looks dangerous or stupid or beautiful, etc.
- generalizations 2y agoClaude has been pretty great. I stood up an 'auto-script-writer' recently, that iteratively sends a python script + prompt + test results to either GPT4 or Claude, takes the output as a script, runs tests on that, and sends those results back for another loop. (Usually took about 10-20 loops to get it right) After "writing" about 5-6 python scripts this way, it became pretty clear that Claude is far, far better - if only because I often ended up using Claude to clean up GPT4's attempts. GPT4 would eventually go off the rails - changing the goal of the script, getting stuck in a local minima with bad outputs, pruning useful functions - Claude stayed on track and reliably produced good output. Makes sense that it's more expensive. Edit: yes, I was definitely making sure to use gpt-4o
- lagniappe 2y agoThat's pretty cool, can I take a look at that? If not, it's okay, just curious.
- generalizations 2y agoIt's just bash + python, and tightly integrated with a specific project I'm working on. i.e. it's ugly and doesn't make sense out of context ¯\_(ツ)_/¯
- lagniappe 2y agoAlright, no worries. Thanks for the reply
- SparkyMcUnicorn 2y agoMy experience reflects this, generally speaking. I've found that GPT-4o is better than Sonnet 3.5 at writing in certain languages like rust, but maybe that's just because I'm better at prompting openai models. Latest example I recently ran was a rust task that went 20 loops without getting a successful compile in sonnet 3.5, but compiled and was correct with gpt-4o on the second loop.
- 2y ago
- creatonez 2y agoNotably, this prompt is making "hallucinations" an officially recognized phenomenon: > If Claude is asked about a very obscure person, object, or topic, i.e. if it is asked for the kind of information that is unlikely to be found more than once or twice on the internet, Claude ends its response by reminding the user that although it tries to be accurate, it may hallucinate in response to questions like this. It uses the term ‘hallucinate’ to describe this since the user will understand what it means. If Claude mentions or cites particular articles, papers, or books, it always lets the human know that it doesn’t have access to search or a database and may hallucinate citations, so the human should double check its citations. Probably for the best that users see the words "Sorry, I hallucinated" every now and then.
- hotstickyballs 2y ago“Hallucination” has been in the training data much earlier than even llms. The easiest way to control this phenomenon is using the “hallucination” tokens, hence the construction of this prompt. I wouldn’t say that this makes things official.
- creatonez 2y ago> The easiest way to control this phenomenon is using the “hallucination” tokens, hence the construction of this prompt. That's what I'm getting at. Hallucinations are well known about, but admitting that you "hallucinated" in a mundane conversation is a rare thing to happen in the training data, so a minimally prompted/pretrained LLM would be more likely to say "Sorry, I misinterpreted" and then not realize just how grave the original mistake was, leading to further errors. Add the word hallucinate and the chatbot is only going to humanize the mistake by saying "I hallucinated", which lets it recover from extreme errors gracefully. Other words, like "confabulation" or "lie", are likely more prone to causing it to have an existential crisis. It's mildly interesting that the same words everyone started using to describe strange LLM glitches also ended up being the best token to feed to make it characterize its own LLM glitches. This newer definition of the word is, of course, now being added to various human dictionaries (such as https://en.wiktionary.org/wiki/hallucinate#Verb https://en.wiktionary.org/wiki/hallucinate#Verb) which will probably strengthen the connection when the base model is trained on newer data.
- devit 2y ago<<Instead, Claude describes and discusses the image just as someone would if they were unable to recognize any of the humans in it>> Why? This seems really dumb.
- ForHackernews 2y ago"When presented with a math problem, logic problem, or other problem benefiting from systematic thinking, Claude thinks through it step by step before giving its final answer." ... do AI makers believe this works? Like do think Claude is a conscious thing that can be instructed to "think through" a problem? All of these prompts (from Anthropic and elsewhere) have a weird level of anthropomorphizing going on. Are AI companies praying to the idols they've made?
- bhelkey 2y agoLLMs predict the next token. Imagine someone said to you, "it takes a musician 10 minutes to play a song, how long will it take for 5 musicians to play? I will work through the problem step by step". What are they more likely to say next? The reasoning behind their answer? Or a number of minutes? People rarely say, "let me describe my reasoning step by step. The answer is 10 minutes".
- cjbillington 2y agoThey believe it works because it does work! "Chain of thought" prompting is a well-established method to get better output from LLMs.
- gdiamos 2y agoWe know that LLMs hallucinate, but we can also remove them. I’d love to see a future generation of a model that doesn’t hallucinate on key facts that are peer and expert reviewed. Like the Wikipedia of LLMs https://arxiv.org/pdf/2406.17642 https://arxiv.org/pdf/2406.17642 That’s a paper we wrote digging into why LLMs hallucinate and how to fix it. It turns out to be a technical problem with how the LLM is trained.
- randomcatuser 2y agointeresting! is there a way to fine tune the trained experts, say, by adding new ones? would be super cool!
- AcerbicZero 2y agoMy big complaint with claude is that it burns up all its credits as fast as possible and then gives up; We'll get about half way through a problem and claude will be trying to rewrite its not very good code for the 8th time without being asked and next thing I know I'm being told I have 3 messages left. Pretty much insta cancelled my subscription. If I was throwing a few hundred API calls at it, every min, ok, sure, do what you gotta do, but the fact that I can burn out the AI credits just by typing a few questions over the course of half a morning is just sad.
- dlandis 2y agoI think more than the specific prompts, I would be interested in how they came up with them. Are these system prompts being continuously refined and improved via some rigorous engineering process with a huge set of test cases, or is this still more of a trial-and-error / seat-of-your-pants approach to figure out what the best prompt is going to be?
- beefnugs 2y ago"oh pretty please? digi-jobs if you are super helpful for internetbucks!" there is no way that they are testing the effectiveness of this garbage
- slibhb 2y agoMakes me wonder what happens if you use this as a prompt for chatgpt.
- deleted 2y ago[deleted]
- sahli 2y ago[dead]