4 ms·
No source code, as usual. How difficult would it be to duplicate these results with TensorFlow? Would something like this require more than the building blocks
by cypher543 10y ago
No source code, as usual.
How difficult would it be to duplicate these results with TensorFlow? Would something like this require more than the building blocks that TF and other toolkits provide? I have zero experience with machine learning, so I'm just curious.
- bbayer 10y agoI think implementation can be done with TensorFlow but hardest part is to provide relevant training set to model.
- yorwba 10y agoThey mention implementing the model in TensorFlow, so it should absolutely be possible. (Unless they are using a special Google-only version.) One thing that might make reproduction difficult is that they don't describe their initialization procedure. But maybe it just doesn't matter.
- ganfortran 10y ago> How difficult Very difficult: 1.Machine learning is pretty data dependent, and make those datasets are very expensive. Google is not likely to give them away for free, because it is their competitive advantage. 2.The infrastructure to train those models are hard to get outside of Google. Pretty sure it is 10s or 100s of GPUs, with Infinity Band connected PS server, running for days and weeks. Even with source code published, people will still have to scratch their head to duplicate Google's performance. Until the day, some equivalent organization as GNU that democratize data access to the public and some mighty algorithm being discovered dramatically reduced the computational requirement for training those models, Google succeeds by just being Google is unlikely going to change.
- cypher543 10y agoRight, even a good diphone voice needs lots of data. And I noticed they trained it with the existing Google Home voice actress, from whom they must already have many, many hours of recordings. I was mostly asking about the model itself; whether you could download TensorFlow and put one together based on this paper alone.
- ganfortran 10y agoI see your points. But it is related. Even if u get what you think the paper describes, it is hard to know whether you did it right or not, because you cannot replicate the result easily. This happens in a lot of CV papers already, where people reimplement the model, but it never get as good as the paper demonstrated But, you have a very good idea. Since it is Google Home, will it be possible that some people just buy hundreds of them, and infinitely ask them question to gather the training data? That will be interesting to see.
- wiz21c 10y agoYou're describing the future : 100% capitalistic society... I'm not sure I like it...
- striking 10y ago1. If you can pay for around 24.6 hours of VA speech data, you can get enough data to run this process with the same quality that Google presented. (that's from the "Experiments section") Not cheap (definitely not free, especially considering the amount of quality control you have to apply), but not expensive either. 2. You can rent out a 96GB GDDR5 GPU instance from Google's cloud for pretty cheap. (https://cloud.google.com/compute/docs/gpus/ https://cloud.google.com/compute/docs/gpus/) I don't think you need anything more powerful than that (but feel free to prove me wrong). I think your last paragraph is totally misguided/uninformed. You can download models for cheap/free (for non-commercial/edu use) from UPenn (https://www.ldc.upenn.edu/language-resources/data/obtaining https://www.ldc.upenn.edu/language-resources/data/obtaining). People don't give away models for free with 0 strings attached because they're a pain to make. And if you want something you can run on a home computer for cheap/free, you can try DeepSpeech: https://github.com/mozilla/DeepSpeech https://github.com/mozilla/DeepSpeech. All you need is an Nvidia GPU.
- Animats 10y agoNo source code, as usual. So why is it on Github?
- speps 10y agohttps://github.com/google/tacotron https://github.com/google/tacotron Still no source code...
- geon 10y agoEasy way to publish?
- pulse7 10y agoBUT: They have a license already there (ASL2). Maybe they will publish the code...
- striking 10y agoTo make the voice samples easily available. Preprint archives and scientific journals don't usually allow you to embed audio into your papers.