3 ms·
Lots of domain experts casually mention that there's nothing defensible in a well trained model. If that's what gets you going: Russia could also match or beat
by some1else 9y ago
Lots of domain experts casually mention that there's nothing defensible in a well trained model.
If that's what gets you going: Russia could also match or beat America in AI.
- kyleschiller 9y agoThe article is subtitled "Its deep pool of data may let it lead in artificial intelligence" for a reason, you can't just pull a "well trained" model out of thin air, and proprietary data is absolutely defensible.
- visarga 9y agoYou'd think China would have an advantage here. But the kind of special data only Google and FB have is mostly related to ads and people profiles. That's not the kind of data you need to advance AI. Instead, what is needed is user labeled images, videos, text, speech and of course, simulations (games, VR) for virtual robotics and reinforcement learning agents. These kinds of data and simulations are open sourced and growing. Once there is a sufficient collection of training data, it's cheap to train your neural nets. Even today, there is much more data than scientists care to use. For example, ImageNet is over 1TB, but a smaller portion of it is used routinely (it has 21841 topics but usually just 1000 of them are used in research papers). So it's not the lack of images that's keeping them. China would have no advantage here. Everyone has approximately the same level, the best paper of last year is the average of this year. That's about AI in general - but regarding self driving cars and ads Google is cooking in secret so it might be a little more ahead.
- cat199 9y ago> But the kind of special data only Google and FB have is mostly related to ads and people profiles. That's not the kind of data you need to advance AI. > Instead, what is needed is user labeled images, videos, text, speech. umm: google image search, all html embedded/related metadata, adwords itself serving as a refined metadata/tagging and ontology database building engine, recaptcha 'how many buildings are in this picture' puzzles, android debug settings where things are uploaded, google drive cloud storage, gmail, the google speaker gizmo and voice activated phone agent, google voice voip, google chat, maps annotations to pictures/gps coordinates, etc etc etc and yadda yadda yadda... not sure what else you need as far as those sources..
- visarga 9y agoEveryone has access to image search results (from multiple providers). Ontologies such as DBpedia and WikiData with billions of triples are open source, don't see why they couldn't closely track closed-doors ontologies such as the one used by Google. The picture puzzles are useless - we already have superhuman ability to classify traffic signs and storefronts. Android debug data and map data are irrelevant for basic AI research.