6 ms·
*When asked by GPT4 to compare the outputs. I'm a staunch believer that it would be foolish to rely on GPT4 for quality comparisons, and it has been mind boggl
by Jackson__ 3y ago
*When asked by GPT4 to compare the outputs.
I'm a staunch believer that it would be foolish to rely on GPT4 for quality comparisons, and it has been mind boggling to see so many people do it and treat it as perfect proof of anything.
It would be slightly more understandable if there was a study to see how human and gpt4 preferences compare, but I'm unaware of any such thing.
- letmevoteplease 3y agoThere is one: "The agreement between GPT-4 and humans reaches 85%, which is even higher than the agreement among humans (81%). This means GPT-4’s judgments closely align with the majority of humans. We also show that GPT-4’s judgments may help humans make better judgments. During our data collection, when a human’s choice deviated from GPT-4, we presented GPT-4’s judgments to humans and ask if they are reasonable. Despite different views, humans deemed GPT-4’s judgments reasonable in 75% of cases and are even willing to change their choices in 34% of cases."[1] [1] https://arxiv.org/abs/2306.05685 https://arxiv.org/abs/2306.05685
- lolinder 3y ago80% agreement is high, but the margins between models at the top are so low that even that remaining 20% could be enough to alter the final rankings, depending on which direction it errs.
- bagels 3y agoNot just that, but are the 80/20 randomly distributed? Probably not. These comparisons might have more in common with the 20%.
- s17n 3y agoBut 80% agreement is the same as between humans
- lolinder 3y agoTrue, but given that you're using one of the models to judge the others, it's likely that the cases of disagreement will tend to favor GPT-4. You would never use one of the competitors as a judge among humans.
- sdenton4 3y agoIf only there were some way to produce some kind of "interval" of scores where you were confident that the actual score sat, and then had some way of comparing these intervals between the different models...
- bostonsre 3y agoIs perfect agreement possible? And what is the definition of agreement? Humans don't agree about much.. are we saying agreement means it matches the truth after intensive investigation by humans?
- shanusmagnus 3y agoUgh, when I was doing my PhD work we were studying creativity in an experiment, and we needed an assessment for how creative different solutions were, and trying to get inter-rater reliability on this quite simple thing was just agonizing. I wound up abandoning the experiment because getting enough reliability would have required screwing down the standards so tightly that it would have ruined the underlying point of the thing.
- biomcgary 3y agoKind of hard to consistently evaluate creativity when someone might pull a James T. Kirk (https://en.wikipedia.org/wiki/Kobayashi_Maru https://en.wikipedia.org/wiki/Kobayashi_Maru), which is only creative the first time and just a cheat thereafter.
- shanusmagnus 3y agoJust one of many problems :)
- drittich 3y agoIt's surprising to me that you describe creativity as a simple thing. How were you defining and measuring it?
- shanusmagnus 3y agoI don't (and didn't) think creativity was simple, but the task was super simple (alternate uses task). The fact that people couldn't agree on how creative the answers were to such a simple task was the insight, although maybe it was only an insight because I was naive.
- sdenton4 3y agoThat's helpful! I've done a lot of work in audio synthesis, which is notoriously difficult measure. The gold-standard is human ratings of audio quality, but it is tough to design good tests (easy to fatigue raters) and the iteration time waiting for results is quite long. Instead, there's now some projects which use neural networks trained on human ratings to predict audio quality, such as ViSQoL: https://github.com/google/visqol https://github.com/google/visqol This opens up fast iteration - scores going up generally corresponds to higher quality - followed by human testing at major milestones (eg, releasing a paper/model). VISQOL has a harder time comparing 'unrelated' models, IMO - ends up being not so great for comparison of different techniques, but excellent for measuring incremental improvement or catching regressions. But, in the end, yes - you can use NN's to measure the quality of other NN's, so long as you're careful about it and make use of human raters from time to time as well. The problem of test data getting into the training data seems to be an especially pernicious issue with LLM's, which isn't really arising in the audio synthesis space.
- blackkettle 3y agoI’ve started doing this with ASR hypotheses from colloquial spontaneous speech. It tends to have similar issues. Lots of shady human ground truth especially where addresses, alphanumeric sequences, repairs and repetitions and other essentially non read speech are concerned. The very large Whisper models are consistent in their transcription style and highly reliable as long as you pick strongly represented languages. And ChatGPT can do a very good job at comparing the linguistic coherence of hypotheses from multiple recognizers. Together these models can annotate, analyze and ingest far more data more consistently than human annotators at this point (at least in the best covered languages). We haven’t quite realized this as a community yet though, because the standard datasets we use for evaluation contain all these human inconsistencies. Wild times.
- 2c2c2c 3y agojust curious, are there any open models doing the opposite of audio synthesis? As in able to generate the stems for a song?
- 3y ago
- koalacola 3y agoIf it's trained by humans is it safe to assume that we'll get it so something crazy like 99% agreeable?
- bottlepalm 3y agoIt's funny how ChatGPT really does give you the most balanced, middle of the road answers. It feels like a distillation of all human knowledge and sentiments. I use it constantly to get advice on plans, architectures, thoughts, etc.. to get an idea of pretty much what the average person would think. It often points out things I've overlooked which I'll improve my design with and go back and forth with ChatGPT until we're both in agreement. I even read a classic book the other day and had a great discussion with ChatGPT about moral relativism, the different schools of thought and how it fit into philosophy as a whole. For students this technology is incredible, I wish I had it for all my classes. Even sometimes comments I'll make on here or Reddit I'll pass through ChatGPT first to see if I made any mistakes in my logic.
- TowerTall 3y ago> to get an idea of pretty much what the average person would think There is no such thing as an average person. https://www.thestar.com/news/insight/when-u-s-air-force-discovered-the-flaw-of-averages/article_e3231734-e5da-5bf5-9496-a34e52d60bd9.html https://www.thestar.com/news/insight/when-u-s-air-force-disc...
- Folcon 3y agoFunnily enough I think you both might be right here, there isn't such a thing as an average person, but ChatGPT may be the synthesis of the average opinion.
- wahnfrieden 3y agowhat is an average opinion? it is the sum of opinions which disagree with the result
- JieJie 3y agoMaybe "balanced" rather than "average" is a better way of putting it?
- reaperducer 3y agoit has been mind boggling to see so many people do it and treat it as perfect proof of anything. The world has long been divided into two camps: People who think computers can make mistakes; and people who think computers never make mistakes, and blame the humans that program them. Well, now the computers are programming themselves. And clearly they're making mistakes.
- deleted 3y ago[deleted]
- courseofaction 3y agoI agree for data creation, but evaluation seems to have little risk of contaminating the outcomes when used with human validation.
- monnow 3y ago[flagged]