5 ms·
"The cost of training the model from scratch using our code is about $50k." Still a substantially steep curve for a bootstrapping startup. It's something I con
by anon1253 7y ago
"The cost of training the model from scratch using our code is about $50k."
Still a substantially steep curve for a bootstrapping startup. It's something I continually run into myself. I have somewhat of a weekend project trying to build a search engine but man ... the cost of just the SSDs and GPUs is daunting on a regular salary. As the complexity of these models grows, so does the barrier to entry for a regular joe like me; which is a shame I think. I know in the US it's fairly normal for a data scientist to pull 100k+ / year, but in the Netherlands salaries pretty much stall at 40k (and angel investment in IT/AI is at an all time low). More generally I fear this will become a bit of a sociotechnical issue if complex AI models will be out of reach for entire economies (especially for cases like language because not everyone speaks English and "minor" languages like those in EU countries are a massive market to explore, yet hard to get into).
- buboard 7y agoThere is no need to train the model, they already provide the parameters. These transformer models , like BERT are pretty adaptable to reuse.
- anon1253 7y agoThere is no need to train this particular model, but adopting this (or any novel model, the field is moving fast) to say Dutch, Italian, Hungarian, Icelandic or whatever still requires training. Luckily for most languages they are provided (at least in the case of BERT, FastText, or regular skipgram). But there is also still quite a bit of leeway in domain specific adoption (for example SciBERT for scientific texts, or legal and financial documents) reddit/wikipedia does carry a bias. Each of which not only requires pretraining the model, but also generating a huge and fairly well formatted corpus. And, although the parameters are usually finetunable, it does break down sometimes on the various sub-word tokenizations used.
- tzapzoor 7y agoAnd they're just "two masters students, with no prior experience in language modeling" with $50k lying around for training a huge model.
- est31 7y agoThe money came from Google: "We would like to thank Google (TensorFlow Research Cloud) for providing the compute for this".
- p1esk 7y agoThey mentioned they spent $500k (in research credits) on all the experiments to actually find the hyperparameters.
- godelski 7y agoWhere did you see that? Also how do two masters students with no experience in NLP get $50-500k in compute credits? How do I get that deal?
- p1esk 7y agoIt was in the article last night, they deleted it for some reason after my comment. One of the authors has a peer reviewed NLP publication [1], the other has several publications in computer vision. I don’t know how they got research credits from Google. [1] https://arxiv.org/abs/1905.13153 https://arxiv.org/abs/1905.13153
- godelski 7y agoWell that is pretty dubious. That completely changes the metric of how much money and how difficult this costs. (And I definitely believe you because there's several other comments with that same number). Could be a typo? But I feel like that's something you'd say. Then again, these are masters students.
- p1esk 7y agoIt's $50k per training run. Their main contribution has been finding the optimal hyperparameters, not described in the OpenAI paper. Obviously you need more than one training run to to that.
- sytelus 7y agoSeriously, you do not need massive amount of money to build search engine. The V0 version of Google was literally a single cheap desktop running whole thing end to end. Average webpage size even today is just 2MB. You can just start with crawling 1B pages which you can comfortably store and index on 8TB HDD (you don't need SSD at this stage). You should make your algorithm work for single user on this 1B page crawl. Don't worry about freshness or speed, just focus on relevance. If things works beautifully, find a VC and show them the goods. If things works out, you might probably end up with $100K seed money. Then you go out and buy dozen good desktops, increase crawl size to may be 15B pages and support few dozen simultaneous users. Now you can go out to big guns, send them a link to your search engine and get next level of funding. At that point you hire real employees and now your job is scaling up to thousands of invite-only users and squeezing out the long tail of performance in terms of relevance. Constantly measure your customer satisfaction and complaints, use gmail-like invite your friends model to incrementally get high-value users who are willing to try out something new. At this point, if your search algo is really much better, you should be able to plot exponential DAU/MAU growth and get funding rounds in $100M range to really scale up.