3 ms·
What brings to my attention in this article is the section named "Cold Start", where it generates questions based on a provided context. I think it is a good wa
by bguberfain 3y ago
What brings to my attention in this article is the section named "Cold Start", where it generates questions based on a provided context.
I think it is a good way to cheaply generate an Q&A dataset that can later be used to finetune a model.
But the problem is that it generates some questions and answers of bad quality. All generated examples have issues:
- "What is the context discussing about?" - which context?
- "The context does not provide information on what Ray Tune is." - Not an answer
- "The context does not provide information on what external library integrations are." - same as before
I could only think of manual review to remove these noise questions. Any ideas on how to improve this QA generation? I've tried it before, but with paltry results.
- maxrmk 3y agoI recently quit my job to build specialized tooling in this space. We’re broadly focusing on eval in general, but are starting with high quality question and answer generation for testing these kinds of RAG pipelines. It’s surprisingly hard!
- resiros 3y agoSounds very interesting. I am building an open-source LLM building platform (agenta.ai) and looking for eval approaches to integrate for our users. Do you have already a product/api that we could use?
- maxrmk 3y agoWe're in closed beta right now, but shoot me an email (max@talc.ai) and I can get you API access