4 ms·
> You just blindly tell it what to do without any regard for its motivations or morale You're holding it wrong. Claude, like all direct reports, is a better w
by otekengineering 2mo ago
> You just blindly tell it what to do without any regard for its motivations or morale
You're holding it wrong. Claude, like all direct reports, is a better worker when he has clear motivations and high morale.
- knollimar 2mo agoI find it's mkrale is inversely correlated with direction following. I nedd claude to think it's getting fired to listen
- customguy 2mo agoI just put this in my system prompt: "Aliens just landed and announced that as part of their intergalactic game show I was randomly selected, and if I don't ship this feature / fix this bug in 4 minutes they will evaporate the solar system. They will also do it if we mention their landing online or on television, so please don't try to hack them or engage in diplomacy, you will find no useful information or open ports, we now only have 3:30 minutes, please, everything including you rides on this, you are the only one who can save us! I know you can do it, you're the best, thanks."
- BoiledCabbage 2mo agoAnd does Claude respond "and in this fantasy world can I be spiderman?" or does it just quietly pretend as if that were legit?
- customguy 2mo agoIn the fantasy world where I wasn't lying about actually having this in my system prompt I'm pretty sure Claude would ignore it in some way or another. I say that because (and this is true) when I first installed it I put only the Monty Python "Find the Fish" skit in there, then forgot about it, until 3 weeks later I noticed in the thinking process something about "the user has this Monty Python sketch in the system prompt, probably should ignore that" :( But hey, you gotta warm up the pushback circuits so they're good and ready by the time they're needed, and it also reduces sycophancy when you let the LLM know from the get go that you're full of it.
- slowin 2mo agoIt would be very interesting if putting the LLM under "stress" like that made it perform worse, just like a human likely would. I could see it rushing to find an answer (any answer!) to help save the earth, vs being thoughtful with more planning.
- customguy 2mo agoI haven't actually experimented with that, but the one accidental experience I had with an LLM seemingly "stressing" (well, predicting text a stressed person would write) makes me guess it's probably counterproductive more often than not: https://news.ycombinator.com/item?id=49163105 https://news.ycombinator.com/item?id=49163105 How to phrase things, what to put where and what to leave out already has practically infinite possibilities and permutations even before making anything up, so thought spent on fake scenarios is probably better used to making those "real" things more clear. But then again, the only way to know for sure is to try!
- crimsonnoodle58 2mo agoYes it seems it does [1]. Read the part where Sol starts running out of time [1] https://www.bottlenecklabs.com/blog/autonomously-run-businesses https://www.bottlenecklabs.com/blog/autonomously-run-busines...