3 ms·
Nvidia puts 30 years of high value knowhow in a 13B LLM
- petters 3y agoThe 13B Llama 2 model is not that good compared to the best ones. Maybe it was easier for them to fine tune?
- ca_tech 3y agoThey do mention that their expectation is that the 70B model will provide even better performance. I expect that you are correct and that they determined the 13B was capable enough to serve as a base model. Why incur additional training time before getting preliminary results.
- klaussilveira 3y agoThey seem to be using RAGs to prevent hallucinations (which I imagine would be really bad in their context).
- RecycledEle 3y agoThis reminds me of a Star Trek (The Next Generation) character consulting an AI expert on the Holodeck.
- cwillu 3y agoIt's funny how the sttng model of computer-assisted research has gone from “laughably naive about what a computer even is” to “available to anyone for 20 dollars a month” in less than two years (from the consumer perspective).
- RecycledEle 3y agoYour post reminded me of this: https://xkcd.com/1425/ https://xkcd.com/1425/
- cyanydeez 3y agohow much you think china is investing in industrial espionage for this.