3 ms·
Takes way more than that in my experience. Back prop isn’t sufficient for rapid generalization.
by valine 3y ago
Takes way more than that in my experience. Back prop isn’t sufficient for rapid generalization.
- byyoung3 3y agoWell, I think you definitely need a certain level of scale, but I think it definitely still works. Also, generalization and memory are two very different things. Generalization is basically iQ, which is very much still a work in progress.
- valine 3y agoIf you have rapid generalization you don’t need scale. Large datasets are only necessary to compensate for the lack of good generalization. The model I posted in my earlier comment responds in character for all queries and was trained in 60 seconds with a dataset smaller than this comment thread.
- byyoung3 3y ago"If you have rapid generalization you don’t need scale" As of right now, you do need scale for rapid generalization. The technology for NLP generalization without scale (both model and data) does not exist. Not saying it won't in the future, just not right now
- valine 3y agoIt does exist, I’ve been playing around with it all week :). Let me know if you’d like a custom Mistral 7B I’ll train one for you. From the output of bing chat, and the very specific way it goes off the rails, I suspect Microsoft has figured it out too. The algorithmic jump to rapid generalization isn’t hard to make. I would be shocked if there’s not an open source version of it a year from now.