4 ms·
I think you are spot on, a LLM focused on science and medicine has a much higher bar to pass when it comes to accuracy. I tried Galactica when it came out, and
by epups 3y ago
I think you are spot on, a LLM focused on science and medicine has a much higher bar to pass when it comes to accuracy.
I tried Galactica when it came out, and I have to say that subjectively at least the results looked much inferior to what the benchmarks suggest. In the paper they claim to be substantially better than GPT-3 on their larger models, while in my personal experience even some generous queries produced garbage output. I cannot remember whether the version available at the site was the largest model, however.
- ipsum2 3y agoGPT-3 was prior to ChatGPT. Since it was a non-instruction tuned model, the output was quite bad unless prompted precisely.
- danielbln 3y agoGPT-3 was instruct finetuned a while before ChatGPT was even released (see InstructGPT paper and announcement). What ChatGPT went through is RLHF to align its responses to a more conversational level, but you could give GPT-3 (davinci) instructions long before ChatGPT.