4 ms·
I'm curious to know what workloads you think Google is running on the same hardware that cloud customers are using.
by honkhonkpants 10y ago
I'm curious to know what workloads you think Google is running on the same hardware that cloud customers are using.
- matt_wulfeck 10y agoAll kinds of CPU intensive work. Isn't that most work?
- hueving 10y agoNo, a lot of work is IO bound. That's why M:N threading (e.g. Go's coroutines, python's gevent) is so popular.
- jfoutz 10y agoWhat an odd question. Seems like training a new voice recognition model, or image classifier, or any of a zillion other research projects google is working on would be perfect for any unused capacity. I've never worked at google. Do they just let each team go with their own crazy hardware setup? Or maybe two separate systems, one internal and one external? That seems a little wasteful (they have the money, so whatever). But wouldn't programming against a big unified api be a bit more, well, sane? I can see really sensitive stuff locked away on its own hardware, but the search front end? why not serve some JS or provide chrome downloads from a giant pool of common hardware? Things grow organically, and they might be technically locked into a specific layout right now, that's totally understandable. But, unused capacity is just capital depreciating steadily away with no gain. Soaking up every spare cycle doing something useful should probably be somewhere in the company goals.
- honkhonkpants 10y agoLike I said, I'm just curious. Google has so much capacity that I'm not sure it would be important for their bottom line to move it between Google and customer workloads in quanta smaller than one physical machine. Your theory also requires Google to have a substantial standby workload that is currently not scheduled, doesn't it?
- jfoutz 10y agoIn my very limited experience, trying to answer research questions can soak up all the computation you can throw at it. Consider something like the traveling salesman problem. The only answer is to check every path. There are heuristics to get good answers, but you never know if they're optimal. You can get an answer on your laptop in an hour. You'll probably get a better answer throwing 1000 hours at your algorithm though, and a better answer throwing 10000 hours. You can do a lot of neat things if p=np. Since no one knows if that's true, we're stuck with bad big O. More computer time is the only way out right now. The point is, I'd be disappointed in google researchers if they didn't have a huge amount of unscheduled workload.
- hueving 10y agoThere are some questions though where the saved energy from not running the batch jobs is worth more than the answer (e.g. searching for the next prime).
- dekhn 10y agoWe (google) ran numerous theory problems on Google's computing platform via the Exacycle program. Peter Norvig convinced me this was a bad idea, but mainly to say that time was better spent on disproving the conjecture using pure math, not search. http://norvig.com/beal.html http://norvig.com/beal.html "But Witold Jarnicki and David Konerding already did that: they wrote a C++ program that built a table of Cz(modp)Cz(modp) up to 5000500050005000 , and, in parallel across thousands of machines, searched for A,BA,B up to 200,000 and x,yx,y up to 5,000, but found no counterexamples. On a smaller scale, Edwin P. Berlin Jr. searched all CzCz up to 10171017 and also found nothing. So I don't think it is worthwhile to continue on that path."
- hueving 10y agoIt's not odd from a security perspective. There is a reason the CIA pays massive amounts of money to Amazon for a completely separate set of hardware from the public. Hypervisors aren't perfect and you might not want customers to be one kvm exploit away from perusing someone's Gmail being processed on the same machine.
- packetslave 10y agoAll of them.