3 ms·
Title: "GPT-4 is getting worse over time, not better" Paper title: "How Is ChatGPT’s Behavior Changing over Time?" Paper Abstract: "GPT-3.5 and GPT-4 are the
by capableweb 3y ago
Title: "GPT-4 is getting worse over time, not better"
Paper title: "How Is ChatGPT’s Behavior Changing over Time?"
Paper Abstract: "GPT-3.5 and GPT-4 are the two most widely used large language model (LLM) services."
When are people gonna realize that GPT-4/3.5 != ChatGPT
As far as I can tell, the paper doesn't explain the methodology either, so hard to know if they're actually using "raw" GPT-4 or GPT-4 via ChatGPT...
I hoped that eventually people would realize they are vastly different, and your experience/results with be vastly different depending on which you use too. But that hope is slowly fading away, and OpenAI isn't exactly seeming to want to help resolve the confusion either.
- simonw 3y agoThey do explain their methodology in some detail in the accompanying GitHub repo: https://github.com/lchen001/LLMDrift https://github.com/lchen001/LLMDrift They seem to have been taking to the API directly and requesting the two different model snapshots. I'm not convinced by their methodology generally. It looks like everything may have been run with temperature 0.1, which I don't think reflects most real-world usage for example.
- deet 3y agoIn looking at the paper's continuous mention of "ChatGPT" and the repo README's statement that "You don't need API keys to get started" .. are we sure they weren't using type of tools to talk to the ChatGPT API (via a session token, etc) vs the OpenAI API? I do agree they talk about the API in the paper a lot but I don't see an exact methods statement that they directly accessed the non-ChatGPT API anywhere, unless I'm missing it,
- simonw 3y agoI think the lack of API key note is because that notebook is the one that renders the charts for the paper.
- peddling-brink 3y agoWhich is better? I assumed they were the same. I’ve been getting ok results with chat-gpt4, might I get better results with the api gpt4?
- capableweb 3y agoDepends on the purpose. I don't think the various parameters (like temperature, top_p) are fully known when it comes to ChatGPT, and neither is the "system prompt" they're using. With the API, you have full control and visibility of those. If you really want to compare "performance"/"quality", you'd have to do so via the API, using known and static parameters and locking the model version. None of which is available via ChatGPT.
- deleted 3y ago[deleted]
- carra 3y agoI guess the titles are also changing over time.