11 ms·
Cloud AutoML: Making AI accessible to every business
- tabeth 9y agoAI without tons of data is, well, overkill. That being said, it seems Google is sharing a few pretrained models, which is nice.
- rasmi 9y agoDisclaimer: I do ML-based work in Google Cloud, but I am not on the AutoML team. The post says there is transfer learning involved, which means in practice you need much less data than you would if creating a classifier from scratch. Of course, more (good) data may yield better results, but it seems one of the goals behind this release is specifically to give custom (your own labels, not just generic object detection) high performance image classification to those who don't have access to Google-scale training sets.
- ska 9y agoTransfer learning is hardly a panacea, however much some would like it to be.
- TheIronYuppie 9y agoDisclosure: I work at Google on Kubeflow Can you say more? I don't think anyone is saying it's magic pixie dust, but it does dramatically reduce the amount of data you need.
- eanzenberg 9y agoIt depends on the domain. It works for images because images are the same in time. It doesn’t work as well for text because there’s tons of nuance to speech patterns between groups (yelp vs google reviews)
- colochef 9y agoyou might want to test www.monkeylearn.com for text
- _delirium 9y agoI'd probably phrase it as "can" dramatically reduce the amount of data you need rather than "does". Getting transfer learning to work in any kind of reliable way is still very much open research, and the systems I've seen are heavily dependent on basically every variable involved: the specific data sets, domains, model architectures, etc., with sometimes pretty puzzling failures. I don't doubt Google has managed to make something useful work, though I'm more skeptical of how general the ML tech is. One advantage of an API like this is that it allows control over many of those variables. I'm not sure if this is what it does, but you could even start out by making a transfer-learning system that's heavily tailored to transfering from one specific fixed model, which coupled with some Google-level engineering/testing resources, could produce much more reliable performance than in the general case.
- ska 9y agoI was about to type a very similar comment, but this is much of what I had in mind. I've also seen it used to justify insufficient validation - resulting in strange generalization failures.
- TheIronYuppie 9y agoDisclosure: I work at Google on Kubeflow As you can see here[1], we do provide quite a bit of information about the accuracy and training of the underlying model. Additionally, the AutoML already (often) provides better than human level performance[2]. Your comment about transferring a heavily tailored model from one model to another is basically what it's doing - it's taking something domain specific (vision) and allowing you to transfer it to your domain. [1] https://youtu.be/GbLQE2C181U?t=1m15s https://youtu.be/GbLQE2C181U?t=1m15s [2] https://static.googleusercontent.com/media/research.google.com/zh-CN//pubs/archive/46180.pdf https://static.googleusercontent.com/media/research.google.c...
- zeroxfe 9y agoUm... I don't think anyone here is saying (or even implying) that.
- petra 9y agoWhat kind of datasets ? what size ? Is it something practical for a small business to gather ?
- eggie5 9y agoFigure 4 of the The Decaf paper shows meaningful learning w/ only 10 examples! https://arxiv.org/abs/1310.1531 https://arxiv.org/abs/1310.1531
- fatjokes 9y agoOn what basis do you make that assertion? Other comments have alluded to transfer learning which can be used to take advantage of small datasets, but independent of neural nets, small datasets can be quite suitable to train some useful SVM/linear regression models.
- tabeth 9y agoI define small as in small business, as in less than 100k records. I also don't consider a linear regression to be "AI". Things like genetic algorithms and neural networks are probably overkill for an average small business. There's also little evidence to my knowledge that transfer learning is really effective on non-vision applications. Besides, the fact that you said that linear regressions can be useful for small datasets is exactly my point.
- bob_theslob646 9y ago>There's also little evidence to my knowledge that transfer learning is really effective on non-vision applications. Source?
- sharemywin 9y agoIf I collect a bunch of data and train a model is the data/model mine or theirs? what if in the future they change their minds and decide to change the TOS? did I just build on top of quicksand?
- TheIronYuppie 9y agoDisclosure: I work at Google on Kubeflow The data remains yours and the model is yours - eg if you delete your account, the data and model goes away (think of it like the data or model you stored on a VM). However, what I think you're looking for is to be able to actually download the model, and I'm afraid that's not possible. Are you looking to avoid lock-in? Or something else?
- sharemywin 9y agooutside audits of data handling practices.
- TheIronYuppie 9y agoDisclosure: I work at Google on Kubeflow I'll make sure we have clear statements/auditing on this! We're already committed to GDPR[1] - is there a specific additional service/certification you'd like to see? [1] https://www.google.com/cloud/security/gdpr/ https://www.google.com/cloud/security/gdpr/
- sharemywin 9y agocurious is that Europe only?
- bistro17 9y agoand addressing model interpretibility?
- chimtim 9y agoThen what is "Auto ML"? It sounds like just another cloud service.
- eanzenberg 9y ago“Making AI accessible to every businees” with image classifiers (because every businees needs image classifiers) ^_^
- cdl 9y agoAnd even if every business did need an image classifier, they would still need a technical person that understands how this AutoML model they train can be integrated into business processes or software applications...
- theDoug 9y agoNot everyone _needs_ accessibility ramps, but we all benefit from them. :) Many businesses need them, but don't have the staff or expertise, many may want them for fun functionality (building their own "Not Hotdog" app), but the aim is ease of model creation for anyone with a bit of data and time. (Disclosure: I work in Google Cloud)
- peatmoss 9y agoI get that it's a metaphor, but I seriously am having a hard time equating automated image classification to accessibility ramps.
- pathseeker 9y agoYou wouldn't if you worked at Google.
- eanzenberg 9y agoBecause google is notoriously bad at profitable products, besides search?
- rasmi 9y ago"Soon, Cloud AutoML will release other services for all other major fields of AI." It's a start! https://cloud.google.com/automl/ https://cloud.google.com/automl/
- cdl 9y agoI think "making AI accessible to every business" is a bit of stretch. While there's no doubt that the AutoML suite will bring tremendous benefits to businesses with recommendation and speech and image recognition needs, it falls short of providing more useful insights such as those gleaned by association rules, clustering (i.e. segmentation), and general probabilistic models. I think that if AI is to be accessible to every business then it will deliver insights rather than the machinery to produce the insights. This is especially true in the context of small businesses.
- geocar 9y ago> it falls short of providing more useful insights such as those gleaned by association rules, clustering (i.e. segmentation), and general probabilistic models. I would also hesitate to build a business relying on Google for those things since I'd likely be competing with Google's actual moneymaker.
- TheIronYuppie 9y agoDisclosure: I work at Google on Kubeflow Can you say more about the specific insights you'd like us to provide? The more specific the better :) Happy to see what we can do!
- cdl 9y agoI'm not personally looking for the insights myself, just an observation from working with SMBs trying to leverage data science more generally to improve their businesses.
- namikaze 9y agoGoogle and others are looking ahead into the future to make AI a commodity. This is just another small step towards that direction.
- cdl 9y agoAnd I'm not discrediting the work or the fact that this is progress... I just don't think that the title of the post is entirely accurate. Mostly with respect to the "every" part.
- jorgemf 9y agoThis has to be very expensive for companies. A good business for Google
- illumin8 9y agoThis seems to be a reaction by Google to the Amazon SageMaker release in November: https://aws.amazon.com/sagemaker/ https://aws.amazon.com/sagemaker/ It's great to see that other cloud providers are acknowledging the talent and training data gaps that many large enterprises face when adopting deep learning. Disclaimer: I work for AWS
- TheIronYuppie 9y agoDisclosure: I work at Google on Kubeflow This is an externalization of the service we use at Google internally called Vizier[1], first discussed publicly in June[2]. The idea is that instead of having to build a model yourself, we can use ML (yes, it uses ML to provide ML) to autotune your model and solve your business problem. Basically, instead of having to deal with all the steps in opening an editor, choosing a algo, tweaking, debugging, etc etc, just provide your structured or unstructured data and we'll help you answer your question (which is what customers actually care about). [1] https://research.google.com/pubs/pub46180.html https://research.google.com/pubs/pub46180.html [2] https://www.youtube.com/watch?v=Z2YL4XJKVpQ https://www.youtube.com/watch?v=Z2YL4XJKVpQ
- illumin8 9y agoSame idea for Sagemaker. Nice to see I get a bunch of instant downvotes - I sometimes wonder why even bother participating in this community.
- rasmi 9y agoI didn't downvote you, but I think the comparison to Sagemaker misses the point. This is literally just uploading labeled data and getting a finely tuned classifier out. Hyperparameter tuning is neat, and both Cloud ML Engine and Sagemaker have that, but (correct me if I'm wrong), only AutoML actually handles all of the model architecture decisions itself using transfer learning and learning2learn. See here for details: https://research.googleblog.com/2017/11/automl-for-large-scale-image.html https://research.googleblog.com/2017/11/automl-for-large-sca... This significantly reduces the level of expertise required to train models, and the AutoML models outperform "expert" human-created architectures.
- prats226 9y agoFor those who don't like to wait for access, https://nanonets.com https://nanonets.com Just upload your training data and we will provide machine learning API automatically.
- alanlewis 9y agoIt's not clear from the demo video, but will this help with labeling data? In my experience, that is the most time consuming part of creating models.
- TheIronYuppie 9y agoDisclosure: I work at Google on Kubeflow Well.... no. BUT, it _does_ support unstructured data so you may not need to label your data at all. As always, YMMV.
- T-A 9y agoFrom https://cloud.google.com/automl/ https://cloud.google.com/automl/ Integration with human labeling For customers with images but no labels yet, we provide a team of in-house human labelers that will review your custom instructions and classify your images accordingly. You will get training data with the same quality and throughput Google gets for its own products, while your data remains private. You can use the human labeled data seamlessly to train a custom model.
- zengid 9y agoI don't think this is 'democratizing' AI but rather centralizing Google's control of a utility service.
- ovi256 9y agoYou would have been right if ML would be as accessible as your hypothetical "utility service" implies. It is not, and getting farther from it. If you compare ML to electricity, we're still in the stages where a few players have found that electrifying their manufacturing plants makes sense. Small players can't afford he investment in machinery and skills. Maybe when the machinery is hidden behind a "utility" provider (which would also bring down the skill level) they will.
- zengid 9y agoWhat makes it inaccessible? Are GPUs prohibitively expensive? Are pretrained models unavailable? Is the software source code closed off?
- eggie5 9y agoI would say training a CNN from scratch or even fine-tuning one takes a lot of domain knowledge and best practices which often are not standardised yet. Besides, we don't even know why they generalise in the first place! See: https://arxiv.org/abs/1611.03530 https://arxiv.org/abs/1611.03530, https://arxiv.org/abs/1711.11561 https://arxiv.org/abs/1711.11561
- frahs 9y agoI think the limit is developer talent. It's hard to find people with the right background to train an AI model.
- whoisjuan 9y agoClarifAI has been doing this for 4 years... I really like their service and they have a fair price. It would be interesting to see how it compares (on quality and pricing).
- jorgemf 9y agoI don't think they have been doing it for 4 years as AutoML was quite recent. The idea maybe was there before but noone publish any paper about it before. Bear in mind that this service creates the architecture of the model for you. I think ClarifyAI has a predefined model that is fine tune with the data, which it is not even similar.
- TuringNYC 9y agoCurious if anyone from the Product or Tech team for AutoML could describe how this differs from MetaMind (sadly, now subsumed into SalesForce.) Richard Socher seems to have achieved some of this in 2015 with MetaMind (...or...perhaps he just had a lot of turkers behind the scene hand-crafting networks to fit data drops...)
- kmax12 9y agoDespite the claim to make AI accessible to every business, this release is fairly limited in that it only applies to images. We will have to see how they extend it going forward. Given the technology it's based on, I'd expect things like text, audio, videos to come next. However, I'm curious if they plan to support structured/relational datasets which are definitely something every business needs. In Kaggle's 2017 State of Data Science [0] survey, data scientists said they spent 65% of their time using relational datasets vs 18% for images. Given that Kaggle is owned by Google, this must be something on their radar. For those data scientists, I maintain an open source library for automated feature engineering called Featuretools (https://github.com/featuretools/featuretools https://github.com/featuretools/featuretools). For people interested in trying it out, we have demos (https://www.featuretools.com/demos https://www.featuretools.com/demos) to help you get started. [0] https://www.kaggle.com/surveys/2017 https://www.kaggle.com/surveys/2017
- jorgemf 9y ago> data scientists said they spent 65% of their time using relational datasets vs 18% for images It will be interesting to see the trend over years. One year doesn't say anything about the trend in the industry.
- TuringNYC 9y ago> data scientists said they spent 65% of their time using relational datasets vs 18% for images Part of the reason is that use cases are driven by limited definitions of ML or "AI" -- AutoML for example note they are "Making AI accessible to every business..." but they are just managing a small part of one type of ML (images with conv nets and res nets.) A senior exec who reads this might develop a narrow view of what AI is. A general annoyance is how obvious techniques like regression, tree models, Bayesian models, etc on tabular data are so ignored while everyone gets hyper-obsessed over GANs or whatever. Almost 90% of low-hanging-fruit I see can be captured with simple classic ML applied to tabular data.
- jorgemf 9y ago
- deleted 9y ago[deleted]
- benkarst 9y agoGoogle wants to monopolize, not democratize, AI. I wonder if Google tests for Doublethink skills before you can get hired there now.
- manigandham 9y agoPerhaps you mean monetize instead? They certainly aren't the only ones who can do ML.
- benkarst 9y agoMonetize and monopolize are synonymous in Silicon Valley. Read Thiel's Zero to One.
- strin 9y agoDemocratization should allow users to "own" their models. This is not the case in Cloud AutoML. Users cannot download their models and host them elsewhere. This dependency means Google can have control over the business's AI capabilities.
- jorgemf 9y agoDoes it really say it somewhere? As far as I know when you train TensorFlow models they are stored in gs, I thought this would be similar. Otherwise I don't know how they are planning to integrate this with the api they have to upload your models and make requests.