5 ms·
This tool doesn’t benchmark based on how a model actually responds to the generated prompts. Instead, it trusts GPT4 to rank prompts simply in terms of how well
by fatso784 3y ago
This tool doesn’t benchmark based on how a model actually responds to the generated prompts. Instead, it trusts GPT4 to rank prompts simply in terms of how well it imagines they will perform head-to-head. Thus, there’s no way to tell if the chosen ‘best prompt’ actually is the best, because there’s no ground truth against actual responses.
Why is this so popular, then (more popular than promptfoo, which I think is a much better tool in the same vein)? AI devs seem enamored with the idea of LLMs evaluating LLMs —everything is ‘auto-‘ this and that. They’re in for a rude awakening. The truth is, there are no shortcuts to evaluating performance in real world applications.
- donkeyboy 3y agoThere is a paper on arxiv saying that GPT4 correlation with human evaluators on a variety of tasks with strongly positive. I am also uncomfortable with it, but using GPT4 as a grader is not as bad as you think.
- fatso784 3y agoYou’re missing the point here. It’s not even getting the LLM’s opinion on evaluating the responses to the prompts (which itself is fraught for some tasks, and benchmarks are known to be limited —even OpenAI admits this, it’s why they made evals). It’s one level abstracted from that. It’s evaluating what the LLM thinks of how well the prompt will do, in purely hypothetical terms. That’s hogwash —different LLMs perform very differently even for the same prompts. Try any tool that lets you compare model responses side-by-side. Unless I see actual use cases, this is yet another iteration of overtrusting AI. Here is what HN was talking about, nearly three months ago -the exact same type of ‘auto-prompt-gen’ tool: https://news.ycombinator.com/item?id=35660751 https://news.ycombinator.com/item?id=35660751
- duskwuff 3y ago> Here is what HN was talking about, nearly three months ago -the exact same type of ‘auto-prompt-gen’ tool. I was reminded of the same thing. What a lot of it boils down to is that LLMs have no innate ability to self-reflect. They can pretend to do it, but no more effectively than an untrained human would.
- ChikkaChiChi 3y ago> They can pretend to do it, but no more effectively than an untrained human would. Which is exactly as much as Generative AI should be trusted.
- jstanley 3y agoThat's definitely begging the question. If you're prepared to accept that GPT-4 can answer questions just as well as humans can, why do you even need to do prompt engineering?
- damascus 3y agoHumans still need 'prompt engineering' to answer questions more accurately though. * What's the best way to get to Radio Shack from here? is not the same as * What's the easiest way to get to Radio Shack from memory when riding a bicycle from here?
- smogcutter 3y agoEasiest way to get to Radio Shack on a bicycle is to ride that bike down to Doc Brown’s house, charge the Delorean up to 1.21 gigawatts, and go back in time.
- QuantumGood 3y agoHumans benefit from good communication too. For example, annual U.S. deaths from medical errors is in the hundreds of thousands. Much of it is due to miscommunication. Is this akin to poor human-to-human prompt engineering? Of course, humans will rush and not attempt better communication, and you can take all the time you wish with an AI. And AI will continue to incorporate better prompt engineering that you won't have to write out. But there will always be a continuum from good to bad for communication, and communication outcomes.
- therein 3y agoYou're forgetting what you may consider to be factual, self-evident and a priori is your opinion. You may be under the impression that annual U.S. deaths from medical errors being in the hundreds of thousands miscommunicates but that is truly your opinion. You are merely jumping to conclusions at places another person might not. And going on to rely on the LLM to validate your perspective is a lossy process. It may not lose your perspective but it loses someone else's and you don't even seem to notice or care.
- LASR 3y agoI think the parent poster is saying that it’s grading the prompts and not the output generated from the prompts. Yeah I agree there. Unless you can check against the output, it’s not really telling much.
- DearAll 3y ago> There is a paper on arxiv saying that GPT4 correlation with human evaluators on a variety of tasks with strongly positive Could you post the link please
- otikik 3y agoOf course there is strong correlation. That is literally what it was designed to do. The problem is that it will simultaneously say that "cow eggs are bigger than chicken eggs", with the same confidence (and in a way that correlates well with human evaluators). https://www.reddit.com/r/Funnymemes/comments/10ohd2n/chatgpt_is_learning_but_still_lacks_real_life/ https://www.reddit.com/r/Funnymemes/comments/10ohd2n/chatgpt... So when you get an evaluation you are playing the russian roulette - you may get a decent result, or you may get cow eggs.
- flangola7 3y agoI just asked and it told me cows are mammals and do not lay eggs. That Reddit post is not even GPT-4 and is 5 months old, which may as well be the 19th century on AI tech timescales.
- albrewer 3y agoThe post you're replying to is a case in point. This time it's cow eggs; what next?
- otikik 3y agoYou are concentrating on the details and avoiding the point. The point is that the tool fails, and it is known to fail, so much that we even have a name for the times when it fails - hallucinations. I have been calling them cow eggs because that's a nice mental image and I didn't want to have to remember for the proper English term. I will continue calling them cow eggs.
- martincmartin 3y agocorrelation ... strongly positive. A positive correlation just means better than chance. "Strongly" is vague, and might not be much better than chance.
- dragonwriter 3y ago> A positive correlation just means better than chance. "Strongly" is vague, and might not be much better than chance. No, adverbs like “strongly” modify adjectives (or verbs, but that’s not relevant here) not nouns; “strongly” is an intensifier that modifies “positive”, its not a separate adjective that modifies the noun “correlation”.
- nomel 3y agoIf we could use GPT-4 to grade prompts, we wouldn't need to talking about grading prompts to use for GPT-4, since this solution requires that the problem doesn't exist. The question then becomes, how do you grade the prompt grading, objectively? At the bottom, there has to be a ground truth. You can't use the thing you're testing to evaluate its own performance. This applies to rulers, speedometers, and AI. It's the difference between a "subjective" and "objective" metrics. If you want an objective metric, you need to have it based on something external, based on reality, objective. Otherwise, you have metrics and ideas that have to held themselves up. Source: My day job is test and measurement. These concepts go back centuries. You never trust your measurement system, you verify it against a standard.
- larodi 3y agoit seems to me also, that this is very much some sort of snake oil for the llm era. prompt generation varies from llm-to-llm and I doubt gpt4 can do reasonable evaluation, provided that it does not know at all about other models.
- londons_explore 3y agoThe various leaderboards show that, in aggregate, LLM's acting as evaluators match other LLM's and humans remarkably closely.
- toxicFork 3y agoWell, you could keep everything else in the project and put yourself or a human as the "does this result feel better than the other one" decision maker
- jasonlotito 3y ago> Why is this so popular Grifters. I won't say that the person working on this is a grifter. Instead, it's so popular right now because of grifters. The same type of NFT grifters and crypto grifters who are mostly silent now. They've moved on. How ethical would it be to sell things to these grifters, to sell the shovels they will use? But I'm always hung up on the idea that they will use those shovels on others and exploit them.
- immibis 3y agoBecause this is the upside if not the peak of the hype bubble. All you have to do is use GPT for a task, and nobody cares whether it actually works, you still get whatever VC funding is left after the interest rate hikes.
- deleted 3y ago[deleted]
- typpo 3y agoThanks for mentioning promptfoo. For anyone else who might prefer deterministic, programmatic evaluation of LLM outputs, I've been building this for evaluating prompts and models: https://github.com/typpo/promptfoo https://github.com/typpo/promptfoo Example asserts include basic string checks, regex, is-json, cosine similarity, etc. (and LLM self-eval is an option if you'd like).
- fatso784 3y agoNo problem! I guess I will make a plug myself --we've been working on a similar 'prompt engineering' tool, ChainForge (https://github.com/ianarawjo/ChainForge https://github.com/ianarawjo/ChainForge). It's targeted towards slightly different users and use cases than promptfoo --geared more towards early-stage, 'quick-and-dirty' explorations of differences between prompts and models for less experienced programmers, versus the kind of continuous benchmarking and verification testing power that promptfoo offers. I particularly like promptfoo's support for CI, which I haven't seen anywhere else, and is very important for developers pushing prompts into production (esp since OpenAI keeps updating their models every few months...).
- mistymountains 3y agoAgreed. This is a pretty terrible idea.