7 ms·
Show HN: TensorFlow for AWS on a Real GPU
- deleted 11y ago[deleted]
- tacos 11y agoTop link points to a non-existent AMI so this "recipe" -- like almost every "here's how to ..." recipe on GitHub that gets posted on HN -- doesn't work. I swear, it's like 100% fail on these types of posts. Also this "real GPU" is explicitly called out in the Google docs as unsupported.
- droque 11y agoI tried with the Oregon region and it failed. I changed to the N. California region and it worked. (No comment about the GPU though.)
- tacos 11y agoGithub recipes are the "works on my machine!" of the new millennium. Just like pecking in programs from computer magazines in the 1980s, only better! I like the ones that are hardcoded to a specific name in a home directory best. Especially when it doesn't match the github name of the "creator."
- fred256 11y agoAn AMI id is region-specific, the original poster should probably have mentioned what region he created the AMI in.
- alexkern 11y agoI've added a note that the AMI's region must be us-west-1. Thanks for the heads up!
- michaelZejoop 11y agoI can't ssh into instance using the AWS documentation (I was pretty careful following the instructions and know I have the right instance id and region)--> "Permission denied (publickey)." (I've since fixed this - I hadn't chmod'd right, and didn't account for working from an ubuntu machine)
- Smerity 11y agoHonestly, there are enough issues with TensorFlow right now due to CUDA 3.0 that using it with AWS is highly problematic. I appreciate the author's attempt, but there's no way the five lines of code he changed to allow CUDA 3.0 has fixed any of the issues found in [1], such as NaNs during training, equally slow training on a g2.2xlarge as a g2.8xlarge, etc ... If you're just interested in playing around, then your laptop will do fine - TensorFlow is happy with just about any hardware you throw at it. Hell, your modern Android phone will run it =] If you're interested in a more involved experiment, develop and debug your task locally on your laptop. By the time you're ready for large scale training, there might be a stable and battle tested AMI such that people are no longer reporting issues in [1] about it. Again, if you're interested, follow the CUDA 3.0 issue on GitHub[1] - this is nowhere near a solved problem and will only cause headaches if you're using it for education. [1]: https://github.com/tensorflow/tensorflow/issues/25 https://github.com/tensorflow/tensorflow/issues/25
- alexkern 11y agoThanks for the feedback! I've added a note to the README that support is still experimental. I'll be tracking the issue and updating the repo + AMI as it develops. Will be compiling with the latest commit (72a5a60) for configurable CUDA Compute support soon.
- vrv 11y agoThanks, would be good to have multiple sources of verification that HEAD now supports this natively without issues like unexpected NaNs. https://github.com/tensorflow/tensorflow/issues/25#issuecomment-156256071 https://github.com/tensorflow/tensorflow/issues/25#issuecomm... is one verification :)
- alexkern 11y agoUpdated to the latest commit, works for me! :)
- erikbern 11y agoI think the nan issue you are referring to was caused by some weird stuff with Bazel. I put together an AMI in virginia: ami-cf5028a5 and if you have a masochistic streak, here are the steps to do it yourself: https://gist.github.com/erikbern/78ba519b97b440e10640 https://gist.github.com/erikbern/78ba519b97b440e10640 The main issue I'm still seeing is that g2.8xlarge with my AMI doesn't run 4x faster than g2.2xlarge even if correctly detects the 4 GPU's. Haven't had time to identify the issue though.
- deleted 11y ago[deleted]