4 ms·
I'm not a scientist but isn't it suspect that they're both creating a new bot and a new evaluation metric for bots at the same time? Like we invented this new
by CSMastermind 3y ago
I'm not a scientist but isn't it suspect that they're both creating a new bot and a new evaluation metric for bots at the same time?
Like we invented this new thing and this new measurement for evaluating it. It does great on the metric we just made up while we were making it.
- htag 3y agoHere is this new thing, and here is how it is different than anything else.
- rjtavares 3y agoWe invented the Turing Test decades ago. Since it became irrelevant with ChatGPT [1], we need new tests. [1]: We can discuss if ChatGPT passes the Turing Test or not, but I think we can now all agree that being able to have a convincing conversation is not a good test for intelligence.
- jpadkins 3y ago[1] I disagree. I think we can agree there needs to be a refinement on the definition of intelligence, but I think LLMs passed the 1950 definition of general machine intelligence.
- deleted 3y ago[deleted]
- theptip 3y agoNo, it’s not suspect in and of itself. Often you need to develop a new benchmark when solving a new problem. It’s common to see this in software engineering/CS papers too. Of course, one should always be critical of benchmarks, and there is an obvious opportunity for bias here that should be reviewed with care. But your phrasing suggests that this is unusual or actively suspicious, which it is not.
- deleted 3y ago[deleted]