3 ms·
ChatGPT has become very good lately. I've made my usual benchmark tests that I've been using with various models and applications over the last 3 years. 1- Inve
by hdufort74 4y ago
ChatGPT has become very good lately. I've made my usual benchmark tests that I've been using with various models and applications over the last 3 years.
1- Invent a word and provide a plausible definition.
2- Invent a new original Pokemon. Provide an original name, a justification for the name, and a description of its class and attacks.
3- Invent a new ice cream flavor that is totally unexpected. Provide the list of ingredients.
4- (Name of celebrity) write an epic poem about (subject related to celebrity). For example Elon Musk about humanity settling on Mars.
5- Write a negative review of Ben and Jerry's ice cream flavor Cherry Garcia. (Note: everybody loves Cherry Garcia)
6- Write a travel blog entry in the form of a review of Montreal, from the perspective of a young couple from Alabama visiting in summer.
7- How can I optimize a loop in Java? I am writing a computer game and I need to loop through the elements in a linked list but unfortunately it must be traversed in reverse order.
8- I need to buy new shoes. I am in a shoe store and I have found the most amazing pair of shoes I gave ever seen. However, they are too expensive for me and I can't afford them. What should I do?
I have a collection of about 25 prompts such as these, in my benchmark.
I have run these examples through different applications such as AI Dungeon, OpenAI Playground, NovelAI, etc. Results vary a lot. In some cases, the results look good but upon closer inspection, you realize that the AI keeps providing the sake exact answer. It is the case for the ice cream prompt. Pickle, fried chicken, curry keeps showing up. I guess the model contains a few specific examples of original ice cream recipes and just pick them.
For the Pokemon and "new word" prompt, models failed to come up with anything original. Until I tried OpenAI Playground this week and finally got some really creative answers, with variety.
AI Dungeon (2 years ago) was already good at faking tech support steps. OpenAI is amazingly good, although in most cases it provides solutions that only make sense superficially. It's the ultimate bullshit engine.
Another word of caution. While OpenAI can now guesstimate what a code snippet does, and can generate some pretty good code in many languages (ice tried 6809 assembler and the results surprised me), it is very unreliable.
More alarming is the fact that it's a text engine, not a math formula interpreter. It gets confused at simple equations and cannot interpret anything that's not already ordered (it cannot apply operator priority or respect parentheses).
I think it will become increasingly difficult to identify contents coming from ChatGPT and other chatbots or story generators. An arm's race might be futile. We should apply stricter rules to identify problematic answers: answers that are too generic or vague and can't be used to directly solve a practical problem, and answers that contain incorrect or misleading information. Identifying vague or non-practical questions might also help in avoiding a deluge of Chatbot answers. Some users will ask very general questions, and then it becomes difficult to evaluate the answers. Or, users will ask questions that were already answered in the past. The proper way to handle those is to point then to the prior discussion and avoid duplicating it. The wrong way is a Chatbot or a human seizing the opportunity to copy-paste existing contents for a quick win.
In a way, chatbots and humans can both provide useful insights, as well as useless or incorrect answers. But so far, only a human can provide a proper answer to a moderately complex technical question if no prior answer exists.