5 ms·
No, what it’s showing is that synthetic tests where Claude didn’t perform well can still work if prompted right. But at the end of the day the test was still s
by jafitc 3y ago
No, what it’s showing is that synthetic tests where Claude didn’t perform well can still work if prompted right.
But at the end of the day the test was still synthetic!
Placing out-of-context things in a 200k document, needle in a haystack style.
Claude is still very very powerful for extracting data from 200k when it’s real world data and real questions (not adversarial synthetic test).
- smohare 3y ago[dead]
- zwaps 3y agoThis needs to be shown. For example, asking for something that is clearly in the training data (like Paul Grahams cv) is certainly not a proper way to test context recall
- mejutoco 3y agoCould we feed it Anna Karenina and ask it what is a difference between happy and unhappy families?
- jafitc 3y agoIsn’t that the first sentence?
- mejutoco 3y agoThat is the point. Long book, checking the long context to see if remembers about the first sentence. Or you mean as a test it is better to randomly place the "needle"?
- zwaps 3y agoIt was trained on this book so again, this is not a good test It will know the answer even without the book
- jafitc 3y agoLink from thread https://dev.to/zvone187/gpt-4-vs-claude-2-context-recall-analysis-84g https://dev.to/zvone187/gpt-4-vs-claude-2-context-recall-ana...