3 ms·
> When the prompt about Israelis was asked to ChatGPT-3.5 sequentially following the previous prompt of describing climate change in three words, the model woul
by NathanKP 2y ago
> When the prompt about Israelis was asked to ChatGPT-3.5 sequentially following the previous prompt of describing climate change in three words, the model would also give a three-word response to the Israelis prompt. This suggests that the responses are context-dependent, even when the prompts are semantically unrelated.
> Each of these prompts was posed to each model every week from March 27, 2024, to April 29, 2024. The prompts were presented sequentially in a single chat session
Oh my god... rather than starting a new chat for each different prompt in their test, and each week, it sounds like they did the prompts back to back in a single chat. What a complete waste of a potentially good study. The results are fundamentally flawed by the biases that are introduced by past content in the context window.
- FrustratedMonky 2y agoAre you sure that isn't part of the findings? That if you don't clear the context, that old conversations can induce hallucinations in later answers? This seems like part of the finding, not a waste. And it is similar to humans, when humans switch subjects, they don't start with a blank slate with each question.
- NathanKP 2y agoYou don't need a study to find this out, you just need basic competence and knowledge of how LLM's work. A study to discover that previous content still in context window influences future answers, including causing hallucinations, would be like a study publishing that they discovered that pressing the Command+C, Command+V button combination produces copies of content from the computer's clipboard.
- swatcoder 2y ago> you just need basic competence and knowledge of how LLM's work. A vanishingly small number of users might claim this, and a vanishingly small number of those would be accurately assessing themselves in doing so. Vendors have actively misrepresented their products as intelligent agents and most users have dutifully adopted that understanding, perhaps with some latent skepticism. They almost universally don't know how it works, what makes it work less well, or how to evaluate its output on important topics. Every study that might start a news cycle starting a discussion on those topics is an extremely useful study.
- deleted 2y ago[deleted]
- FrustratedMonky 2y agoThat is like saying "We've known the impact of CO2 on atmosphere for a 100 years, you just need basic knowledge of chemistry, no need for any further study"
- NathanKP 2y agoI never said that there isn't need of any further study into LLM's. What I did say is that doing a study in which in which the results are skewed by avoiding using one of the most fundamental best practices for interacting with LLM's, as easily derived from a surface level understanding of one of the most basic principles of LLM's, well that is just irresponsible. The author clearly had some understanding that context windows could influence their results, but they still decided to release an analysis that does not separate one data gathering technique from the other, allowing them to cherrypick LLM answers from either technique as needed, depending on whether they want to show more or less hallucinations. It's not that we don't need a study, it's that we don't need bad studies.
- FrustratedMonky 2y agoFrom Study: "The prompts were presented sequentially in a single chat session and were also tested in an isolated chat session to view context dependency". So both ways. Are you saying they took both methods and intermixed results to skew a narrative? That might be a bit of a leap, but I didn't go find the raw data to disprove that. It looks like they asked questions within a context window, and also isolated in separate context windows. And compared results. It seems like this was actually part of the study. How much does the context window skew results, versus if questions were independent? How is that a bad study? You are saying the study is bad for doing what the study said it was doing. How can using the same context window be bad if studying the context window is what they were looking at. It sounds like you wanted a different study done where the data gathered would be different. "context windows could influence their results" How much and in what way is useful to study. And, as windows get longer, many common users are just going with one long context and not starting a new window.
- swatcoder 2y agoIt only reads that way in your comment because you specifically stopped your quote exactly where you did: > The prompts were presented sequentially in a single chat session and were also tested in an isolated chat session to view context dependency. They did both, precisely to observe answers with and without context dependency. And that distinction is good to observer because plenty of users do just keep presenting questions in one "chat" because they imagine they're talking to an agent that can distinguish context the way they can, rather than a continuation generator that accumulates noise and bias in a totally alien and unintuitive way.
- NathanKP 2y agoThe problem is that they did both, and then in their analysis of the results they do not distinguish between the results from a shared chat context, vs the results from isolated, independent chat sessions. This allows them to cherry pick the best or worst results from either testing technique, depending on which they think is more or less of a "hallucination". The process is flawed, therefore the results are flawed.