4 ms·
Prompts aren’t Real
- esafak 8d agoCrafting a good prompt is what differentiates an expert's output from a beginner's. Of course you need good constraints too. But those good constraints are created precisely through good prompts.
- jdlshore 8d agoThis is an amazing article. The problems it describes are exactly what we found when building a production system that used LLMs to (most of the time) produce reliable results. Extensive tests are necessary, and stakeholders have no idea how their suggestions fail in production. They just see the handful of times they tried something and had it work, not the long tail of cursed results. (“How hard can it be? Why don’t you just…”) We didn’t get to the point of self-built prompts, as the article suggests, but it’s an intriguing idea.
- mikepalmer 7d agoThis guy is on the ball with the problem. Totally correct: Like my friend Coda says all the time: "the textual nature of prompts leads us to take the intentional stance towards systems which aren’t conscious, and thus miss the essential nature of their non-meaning." I don't know if his solution (""We should all go insane building interlocking evaluation and optimization pipelines, instead.") would be the long-term solution. Instead perhaps something could be trained into the models, i.e., he is describing a process at inference time that could be done at training time. To make their weird errors less frequent / make them more human. cf. https://arxiv.org/abs/2008.04071 https://arxiv.org/abs/2008.04071 "On Controllability of AI" However, as I said, you can't make it perfect but you can make it better. (You can't make humans fully aligned with human society's interest anyway, including the humans controlling the nukes.)
- ianjbutler 7d ago> I don't know if his solution (""We should all go insane building interlocking evaluation and optimization pipelines, instead.") would be the long-term solution. TFA could explain this one part better I think. The whole process proposed is real with lots of stuff in the literature, but by definition NOT a long-term solution in the sense that this process actually has no end. None of the approaches can get you a static answer for a moving target/platform. So the "interlocking pipelines" for eval/opt would not be some stepping stone you can throw away, and they aren't something you'd run periodically. They'd basically be always on forever and spending 10-100x on system complexity and on tokens. Unless of course you're ready to freeze everything else about the whole system forever (including the backend model, and the whole nature of the "average" context window, the plugins/other prompts in the mix, etc). Are most people in position to freeze requirements/platform forever? Not really, because if they were they'd just build a fairly static system and probably have limited use for AI. Are most people in a position to just casually accept 100x complexity/cost? Not really, that's the "it's not yet webscale" kind of advice that sounds good but isn't necessarily reasonable for average use-case or average org. Since specializing your own locale for this is usually a mistake.. the likely future direction is eval/optimization as a service
- cortesoft 8d agoI hope the future of AI isn't this sort, where companies provide the user/customer with an interface to an AI that can do things for the user... I would much prefer that companies instead provide an interface FOR an AI, and the user brings their own AI which connects to that interface. In other words, provide my AI with tools, instead of providing me an AI that uses your tools. That way, my AI can bring all the context it needs, and I can bring all of the settings and knowledge about what I want with me. I don't want a fractured world of tons of AIs i interact with where I have to explain all the fundamental information about what I want and how I work every time. This also has the benefit of sidestepping the issue the essay is talking about. You provide a consistent tool, and the AI weirdness is not your issue anymore. You don't have to worry about solving for all the weird ways people prompt the AI, or the ways they break.
- kennywinker 8d agoI don’t think that will happen. It sounds good, and I would like it if things operated that way - but from the company perspective how do they, for example, have a tool call that gives the customer a discount, without it getting used when it shouldn’t? Companies want ai to replace human customer service decision making, which means it can’t just be an api that an external agent can interact with, because it needs private knowledge of company processes and access to capabilities that are abusable. But we’re already at the point where if you manage to talk to a human, mostly you end up speaking to someone with no actual power to resolve your issue - so i think basically the future is just going to suck
- visarga 8d agoThey can want anything they like, if customers want to use agents and they don't provide APIs they will lose out.
- adrianN 8d agoYou prevent abuse of the API by models the same way you prevent abuse by humans: you have server side checks.
- justonenote 8d ago
- Joker_vD 8d agoTL;DR: you need to do... essentially supervised learning on your prompts? I mean, if I wanted to do ML, I'd already have been doing it ten years ago.
- mohd_rafay 8d ago[flagged]
- xhxjxchjcdhcxf 8d ago[flagged]
- vouwfietsman 8d agoyour loss
- pell 8d agoFrom the HN guidelines: > Be kind. Don't be snarky. Converse curiously; don't cross-examine. Edit out swipes. >Comments should get more thoughtful and substantive, not less, as a topic gets more divisive. >When disagreeing, please reply to the argument instead of calling names. "That is idiotic; 1 + 1 is 2, not 3" can be shortened to "1 + 1 is 2, not 3." >Don't be curmudgeonly. Thoughtful criticism is fine, but please don't be rigidly or generically negative.
- xhxjxchjcdhcxf 8d ago[flagged]
- visarga 8d ago> the textual nature of prompts leads us to take the intentional stance towards systems which aren’t conscious, and thus miss the essential nature of their non-meaning I see LLMs as being capable of making useful distinctions and having a rich action space. They are widely used because their operation is useful, and that can only happen when semantics work well in practice. But useful things that pay for themselves don't need our "essential nature" blessing, they already have persistence by mutual entanglement with us.
- vouwfietsman 8d agoNo idea how effective this is, but it sure looks a lot more like engineering than most of the 'prompt engineering' things I've seen in the past years. Kudos to the author for writing this concisely without aggrandizing his work.
- roughly 8d agoOne issue with this that we ran into is that it costs actual countable money to run the test suite, which is distinct from anything else I’m used to, so the notion that we’d do enough testing to generate a statistically significant gauge of performance - man, I know it’s correct, but I’m not sure my company will survive the process.
- dvogel 8d agoFor these tests, why not tune the temperature and such to reduce the randomness and convert them to almost-always-succeeds vs almost-always-fails? Is it not the iteration count that drives up the cost?
- sarchertech 8d agoIf turn the temperature down for tests, they won’t match production behaviors. If you turn it down too far in production, the output will just be bad.
- dvogel 8d agoIsn't the goal of the author to get reproducible behavior out of the agent though? I would thinking turning the temperature down would serve that production goal too.
- roughly 8d agoright, this is the whole problem - the stochastic behavior is both the goal and the problem. If you want your tests to match production, you need to get a reasonable sample size, which costs real money.
- sarchertech 7d ago2 things. 1. If you turn the temperature down too far, the output is just bad and no amount of running prompts optimization will let you hill climb your way to good performance. 2. It’s not about determinism vs non-determinism. It’s about chaos. A perfectly deterministic model is still chaotic. Meaning that very small changes to the input result in very large changes to the output. Turning temperature down doesn’t actually get you predictable or reproducible behavior across different inputs.
- dist-epoch 8d agoThe format of this article makes it almost impossible to read. I gave up after about 10 "pages".
- operator3 8d ago[flagged]
- stickr 8d agoAs awesome as this is, because as much as I want to create an agent for my customer that is predictable and "deterministic", I can't help but think of how wasteful, expensive and not-fun this is. Good for the author that it's fun for them, but for me it seems like I am in that "monkey ladder banana" experimemt: doing something because others are and the customer is giving me bana... sorry, money for it, convinced it will help him (the money would 100% stop if I started looping prompt optimizations like this). If I have so many tools and MCPs as I do currently, and with each the behavior regresses and changes wildly, it seems I should either merge tools and do more automations and come back to the prompt. (The alternative being training my own model?)
- rzzzt 8d agoRLHF might help, you don't have to train a model from scratch.
- TeMPOraL 8d agoAm I the only one who waited to the end for, and was disappointed not to see, a peek into these optimized prompts? I so want to take a peek into that abyss, even if that risks the abyss looking back at me. I'm curious just how twisted they get relative to the original, in what alien ways.
- nomad-linkd-id 8d ago[flagged]
- deleted 8d ago[deleted]
- mmargenot 8d agoThis was great! When you think about optimizing prompts with GEPA (or comparable methods and tools), do you consider each tool or skill separately? How do you think about the optimization of the system prompt for a large agentic system? I imagine that you do a collection of passes to cover each overlapping set of what you want evaluated, but the system prompt makes all cases dependent on each other. What I’ve done in the past is use the system prompt to extract subjective criteria for an LLM judge (like various system prompt statements that contribute to brand voice) and check individual traces with that for evaluation, but I’d like to move beyond including that in a prompt at all.
- yt1998 7d ago[dead]
- shshsjsj 8d ago[dead]
- huflungdung 7d ago[dead]
- pjm331 7d ago> That is currently working, but since the fix is fully deranged I expect it’ll be disturbed again at some point. This one made me laugh out loud