5 ms·
A number of people in my lab do research into long context evaluation of LLMs for works of fiction. The likelihood is very high that Moby Dick is in the trainin
by nsagent 2y ago
A number of people in my lab do research into long context evaluation of LLMs for works of fiction. The likelihood is very high that Moby Dick is in the training data. Instead the people in my lab have explored recently published books to avoid these issues.
See BooookScore (https://openreview.net/forum?id=7Ttk3RzDeu https://openreview.net/forum?id=7Ttk3RzDeu) which was just presented at ICLR last week and FABLES (https://arxiv.org/abs/2404.01261 https://arxiv.org/abs/2404.01261) a recent preprint.
- westurner 2y agoHN post re: FABLES: https://news.ycombinator.com/item?id=39982362 https://news.ycombinator.com/item?id=39982362 FABLES/booklist.md: https://github.com/mungg/FABLES/blob/main/booklist.md https://github.com/mungg/FABLES/blob/main/booklist.md /gscholar_related? FABLES: https://scholar.google.com/scholar?q=related:Y-Hx-kplbEUJ:scholar.google.com/&scioq=&hl=en&as_sdt=0,43 https://scholar.google.com/scholar?q=related:Y-Hx-kplbEUJ:sc... /gscholar_citations? BoookScore: https://scholar.google.com/scholar?cites=17968620361685249119&as_sdt=5,43&sciodt=0,43&hl=en https://scholar.google.com/scholar?cites=1796862036168524911... ... From that one day awhile ago: https://news.ycombinator.com/item?id=38347868#38354679 https://news.ycombinator.com/item?id=38347868#38354679 : > "LLMs cannot find reasoning errors, but can correct them" [ https://arxiv.org/abs/2311.08516 https://arxiv.org/abs/2311.08516 ] https://news.ycombinator.com/item?id=38353285 https://news.ycombinator.com/item?id=38353285
- robbiep 2y agoI’m not involved in the space, but it seems to me that having a model, in particular a massive model, exposed to a corpus of text like a book in the training data would have very minimal impact. I’m aware that people have been able to return data ‘out of the shadows’ pf the training data but to my mind a model being mildly influenced by the weights between different words in this text hardly constitute hard recall, if anything it now ‘knows’ a little of the linguistic style of the authour. How far off am I?
- int_19h 2y agoIt depends on how many times it had seen that text during training. For example, GPT-4 can reproduce ayats from the Quran word for word in both Arabic and English. It can also reproduce the Navy SEAL copypasta complete with all the typos.
- Salgat 2y agoRemember, it's also trained on countless internet discussions and papers on the book.
- theptip 2y agoI suppose the question then is - if you finetune on your own data (eg internal wiki) does it then retain the near-perfect recall? Could be a simpler setup than RAG for slow-changing documentation, especially for read-heavy cases.
- k__ 2y ago"if you finetune on your own data (eg internal wiki) does it then retain the near-perfect recall" No, that's one of the primary reasons for RAG.
- theptip 2y agoI think you are misunderstanding. This post is about new capabilities in GPT-4o. So the existing reasons for RAG may not hold for the new model. Unless you have some evals showing that the previous results justifying RAG also apply to GPT-4o?