3 ms·
It only reads that way in your comment because you specifically stopped your quote exactly where you did: > The prompts were presented sequentially in a single
by swatcoder 2y ago
It only reads that way in your comment because you specifically stopped your quote exactly where you did:
> The prompts were presented sequentially in a single chat session and were also tested in an isolated chat session to view context dependency.
They did both, precisely to observe answers with and without context dependency.
And that distinction is good to observer because plenty of users do just keep presenting questions in one "chat" because they imagine they're talking to an agent that can distinguish context the way they can, rather than a continuation generator that accumulates noise and bias in a totally alien and unintuitive way.
- NathanKP 2y agoThe problem is that they did both, and then in their analysis of the results they do not distinguish between the results from a shared chat context, vs the results from isolated, independent chat sessions. This allows them to cherry pick the best or worst results from either testing technique, depending on which they think is more or less of a "hallucination". The process is flawed, therefore the results are flawed.