3 ms·
(I'm a researcher in ML/DL.) > Anyone know why they used CPUs instead of GPUs/TPUs? Cost and resource availability. They're distributing each computation on
by fnbr 7y ago
(I'm a researcher in ML/DL.)
> Anyone know why they used CPUs instead of GPUs/TPUs?
Cost and resource availability.
They're distributing each computation on a different CPU, not distributing each computation over multiple CPUs.
It would be faster to have each computation run on a CPU + GPU, but that would be very very expensive, and hard to schedule.
GPUs/TPUs are also only faster for sufficiently large networks and sufficiently large batch sizes. There's a large fixed cost to send data to/from the CPU, and for smaller networks, it's often not worth running it on a GPU. No idea if this was the case here.
- yazr 7y agoDOes this 2500-cpu-hr cover the ENTIRE learning process? Lets even your first run is crap, and you try again with RGB instead of YUV or whatever. So you do 4 runs. So 10000-cpu-hours replace a week of work of a qualified ML engineer. This is pretty amazing. If i understand correctly.
- fnbr 7y agoI'm not sure about the exact details, but that's my understanding too. It is very exciting. This is directly replacing a week of work (if not more!) of a qualified ML engineer. Additionally, as all of the experiments are cheap (being run on a single CPU) you can run them on the cheap, interruptible, cloud instances, and it's not the end of the world if you need to restart some experiments. This is also not mature science- there's still a lot of active research being done on AutoML- so there's still a lot of potential for improvement.
- so_tired 7y agoIs there any way to "scale up" the solution network for more dimensions? For example, if i train on a a 100x100 image domain, with good results. So now I get a bigger budget to work with 200x200 images. There is no real way to leverage the good architecture from the 1st network. Is there ? Can this be done as a ugly-hack and then be used as a seed into the architecture-search ?
- fnbr 7y agoNot that I'm aware of. I could imagine a few things that you could try that might accomplish this, but I'm not aware of any published literature discussing the efficacy of such work. You could do the naive thing, which would be to take your architecture and scale each layer size up (e.g. select an architecture on Cifar-10 and then scale it up to work on ImageNet). This is done in practice quite often, and seems to work well, but I'm not aware of any robust research done to validate the effectiveness of this.