4 ms·
I'm not surprised at the defence of "prompt engineering" here. It's something easy to do with no real knowledge, and I'm sure having it dismissed hurts some peo
by tensor 3y ago
I'm not surprised at the defence of "prompt engineering" here. It's something easy to do with no real knowledge, and I'm sure having it dismissed hurts some people.
But I 100% agree with the author, "prompt engineering" is not science, and I'd say it's not engineering either. All you're doing is exploring the parameter space of particular model in a very crude way. There is no "engineering" going on in this process, just a bunch of trial and error. Perhaps it should be called "prompt guessing."
None of the results of this process will transfer to any other model. It's simply not science. Papers like "step-by-step" are different, and relate more to learning and inference and do translate to different models and even different architectures.
Also, no, we are not all beginners here. Language models have a long history, and while the very large models are impressive, most of their failings have been known for a very long time already. Things like "prompt engineering" will eventually end up in the same graveyard as "keyword engineers" of the past.
- adastra22 3y ago> But I 100% agree with the author, "prompt engineering" is not science, and I'd say it's not engineering either. All you're doing is exploring the parameter space of particular model in a very crude way. There is no "engineering" going on in this process, just a bunch of trial and error. I wonder what your definition of “science” or “engineering” is…
- 0xfae 3y agoRight? I'm having a hard time imagining a definition that includes "trying new things and seeing what happens" but that doesn't include... "trying new things and seeing what happens"
- aeternum 3y ago"Science" has been twisted recently into a kind of witchcraft that can only be practiced by those anointed through the rigors of academia. "Trust the science" In reality, that is about the furthest from what you should do. As Feynman once said: "Science is the belief in the ignorance of experts". Electricity was also once considered a toy and good for nothing but parlor tricks.
- roguas 3y agoEspecially given this would be fine definition of engineering+science: "All you're doing is exploring the parameter space of particular model in a very crude way."
- gwervc 3y agoIf you remove the AI glasses, "prompt engineering" is just typing words and seeing if results match the expectations... which is exactly what any search engine pays their testers for. Those testers are making an important job to keep improving the quality of the product but they aren't engineers and even less so researchers. Similarly a kid playing with the dose of water needed to build a sandcastle isn't a civil engineer nor an environmental researcher. Maybe on LinkedIn though.
- blueboo 3y agoI’m not sure the scientific method itself can withstand this sort of scrutiny. After all, it’s just making guesses about what will happen and then seeing what happens!
- j2kun 3y agoExcept there's also, you know, building coherent theories and using those theories to predict the system behavior.
- gowld 3y agoThat is in no way a requirement for doing science.
- selfhoster11 3y agoAll right, here is a theory: LLMs contain "latent knowledge" that is sometimes used by the model during inference, and sometimes it isn't. One way to "engage" these internal representations is to include keywords or patterns of text that make that latent knowledge more likely to "activate". Say, if you want to ask about palm trees, include a paragraph talking about a species of a palm tree (no matter whether it contains any information pertaining to the actual query, so long it's "thematically" right) to make a higher quality completion more likely. It might not be the actual truth or what's going on inside the model. But it works quite consistently when applied to prompt engineering, and produces visibly improved results.
- naremu 3y agoIt's funny how often I see people make bring up the "did you know" tidbit about software engineering not being "real" engineering in a traditional sense, which seems to go very uncontroversially. But prompt engineering is still a pressure point for some people, despite being wildly more simple and accessible (literally tell the thing to do a thing, and if it doesn't do the thing right, reword) It feels as though we're getting to the technological equivalent of "what IS art anyways", and questions like if non traditional forms like video games are art (I'm thinking all the way up the chain to even say, Madden games) And in my experience, when something is under constant questioning of whether or not it even counts as X, Y or Z, it usually can technically qualify, but... If people are constantly debating whether or not it's even X, it's probably just not impressing people who don't engage in it, as opposed to "traditional" concepts of engineering and art, and part of the impression made comes from the investment and irreplaceable skillsets, things few, if anyone else at the time could have done. This is why taping a banana on the wall is definitely technically art, but not many outside the art community that tapes bananas to walls really think much of it. It's so mundane and accessible a feat that it doesn't garner much merit to passerbys. It's art by the loosest technical definition, and is giving a lot of credit for a small amount of effort anyone could've done. Admittedly "prompt engineering" is definitely less accessible than a roll of duct tape and a banana but I think we used to just call it "writing/communication", but I guess those who feel capable at that, often just do it manually anyways.
- Tarq0n 3y agoScience should aim to create general (that is, generalized or generalizable) knowledge. One prompt is just an anecdote, a method for creating performant prompts or deriving prompts from model characteristics would be more scientific.
- selfhoster11 3y ago> All you're doing is exploring the parameter space of particular model in a very crude way. Yes. I come to think of prompt engineering as, in a sense, doing an approximate SELECT query on the latent behavioural space (excuse my lack of proper terminology, my background in ML is pretty thin) that can be thought of as "fishing out" the agent/personality/simulator that is most likely to give you the kind of answer you want. Of course a prompt is a very crude way to explore this space, but to me this is a consequence of extremely poor tooling. For one, llama.cpp now has negative prompts, while the GPT-4 API will probably never have them. So we make-do with the interface available. > There is no "engineering" going on in this process, just a bunch of trial and error. Perhaps it should be called "prompt guessing." That is incorrect. It is true that there is a lot of trial and error, yes. But it's not true that it's pure guessing either. While my approach can be best described as a systematic variant of vibe-driven development, at its core it's quite similar to genetic programming. The prompt is mutable, and it's efficacy is possible to evaluate at least in a qualitative sense vs the last version of the prompt. By iterative mutation (rephrasing, restructuring/refactoring the whole prompt, changing out synonyms, adding or removing formatting, adding or removing instructions and contextual information), it is possible to iterate from a terrible initial prompt to a much more elaborate prompt that gets you 90-97% of the way towards nearly exactly what you want to do, by combining the addition of new techniques with subjective judgement on how to proceed (which is incidentally not too different from some strains of classical programming). On GPT-4, at least. > None of the results of this process will transfer to any other model. Is that so? Yes, models are somewhat idiosyncratic, and you cannot just drag and drop the same prompt between them. But, in my admittedly limited experience of cross-model prompt engineering, I have found that techniques which helped me to achieve better results with the untuned GPT-3 base model, also helped me greatly with the 7B Llama 1 models. I hypothesise that (in the absence of muddling factors like RLHF-induced censorship of model output), similarly sized models should perform similarly on similar (not necessarily identical) queries. For the time being, this hypothesis is impossible to test because the only realistic peer to GPT-4 (i. e. Claude) is lobotomised to the extent where I would outright pay a premium to not have to use it. I have more to say on this, but won't unless you ask in the interests of brevity. > Language models have a long history, and while the very large models are impressive, most of their failings have been known for a very long time already. Things like "prompt engineering" will eventually end up in the same graveyard as "keyword engineers" of the past. Language models have a long history, but a Markov chain can hardly be asked to create a basic Python client for a novel online API. I will also dispute the assertion that we know the "failings" of large language models. Several times now, previously "impossible" tasks have been proven eminently possible by further research and/or re-testing on improved models (better-trained, larger, novel fine-tuning techniques, etc). I am far from being on the LLM hype train, or saying they can do everything that optimists hope they can do. All I'm saying, is that the academia is doing itself a disservice by not looking at the field as something to be explored with no preconceptions, positive or negative.