4 ms·
> Code generation: the change they report is that the newer GPT-4 adds non-code text to its output. They don't evaluate the correctness of the code. They merely
by tanaros 3y ago
> Code generation: the change they report is that the newer GPT-4 adds non-code text to its output. They don't evaluate the correctness of the code. They merely check if the code is directly executable. So the newer model's attempt to be more helpful counted against it.
In the prompt they specifically request only the Python code, no other output. An “attempt to be helpful” that directly contradicts the user’s request seems like it should count against it.
- Filligree 3y agoI guess, but that still isn't the sort of degradation people have been talking about. It's not a useful data point in that regard.
- gwd 3y agoI mean, if you were hoping to use the API to generate something machine-parse-albe, and that used to work, but it doesn't any more, then sure, that's a sort of regression. But it's not a regression in coding, but a regression in following specific kinds of directions. I certainly have found quirks like this; for instance, for a while I was asking it questions about Chinese grammar; but I wanted it only to use Chinese characters, and not to use pinyin. I tried all sorts of prompt variations to get it not to output pinyin, but was unsuccessful, and in the end gave up. But I think that's a very different class of failure than "Can't output correct code in the first place".
- oefnak 3y agoThat's false. If it outputs formatted code, it's easier to read. I don't see the backtics, I see formatted code when using the chat interface.