6 ms·
The headline feels like clickbait. Does MLPerf also benchmark power consumption/cost per submission? It's not particularly illuminating that if you use a 10k GP
by joshvm 3y ago
The headline feels like clickbait. Does MLPerf also benchmark power consumption/cost per submission? It's not particularly illuminating that if you use a 10k GPU cluster, you can train things quickly and a "per chip" comparison is vague.
For most people the bottom line is cost and availability anyway - does it matter if a TPU is twice as slow if it costs half as much? I think I read somewhere that Apple Silicon is actually one of the most efficient platforms to use for development (e.g. a tricked out Mini makes quite a good inference server); I've been able to train recent object detection models on my M2 without much trouble (and it's much cheaper than paying for cloud time).
- sillysaurusx 3y agoNah, MLPerf is legit. I was skeptical of it too, but wanted to leave a quick note (I have to run) that they’re solid. I’d be the first to call it out if it was pointless or misleading, but it turns out to be the only way to get a true idea of comparisons across different hardware. It forces everyone to achieve the same goal, which is key; otherwise you’re left with a bunch of slight (and large) variations that tell you nothing about achieving the actual goal, which is all you care about. It’s a bit like a shared space race. Getting that top-N accuracy to 74 point yada yada percent in 37 seconds on a certain resnet architecture tells you that you can do the same thing on that hardware, which means you can do your own things just as quickly. So it’s a worthwhile time investment. If you force everyone to get that same accuracy with the same model arch, you can make informed decisions, which is especially important when throwing millions of VC dollars around. EDIT: Sorry, I completely misread you. It’s difficult to quantify what you’d like to measure: the bottom line price of who is cheaper for your expected workload. There are immense trade offs, not least of which is time spent learning a particular (esoteric) stack. CUDA knowledge doesn’t port to JAX and vice-versa. Prices are always changing, and you can usually work out a deal with the sales team to get lower than advertised. Especially if they want you to choose them instead of some competitor. So in general it’s hard to figure out what you can expect for production workloads in terms of total dev cost vs price vs speed. I will say that as a researcher, there’s no substitute for fast iteration cycles. I’m one of the few who believe in scaling down your models as much as possible when testing experimental ideas, precisely because you can try 30 runs instead of 3. So all else being equal, I’d take speed. But all else isn’t equal. The only thing I want nowadays is free plus stable. It’s looking like a 4090 might be the way to get that, which is enough to try out some interesting ideas.
- Permit 3y agoThe person you’re replying to is not questioning MLPerf, but rather this article’s interpretation of the results.
- sillysaurusx 3y agoOops. Thank you.
- sdenton4 3y ago/CUDA knowledge doesn’t port to JAX and vice-versa./ These are completely orthogonal? Jax can execute on GPU, TPU, or CPU pretty seamlessly, by design.
- sillysaurusx 3y agoAnd then you give up specialized cuda kernels, which is necessary to run mistral 7b in 4 bit float mode. There were some experimental cuda kernel for jax codebases. They didn’t work very well at the time, but maybe it’s better now.
- sdenton4 3y agoI'm still not totally sure what the issue is. Jax uses program transformations to compile programs to run on a variety of hardware, for example, using XLA for TPUs. It can also run cuda ops for Nvidia gpus without issue: https://jax.readthedocs.io/en/latest/installation.html https://jax.readthedocs.io/en/latest/installation.html There is also support for custom cpp and cuda ops if that's what is needed: https://jax.readthedocs.io/en/latest/Custom_Operation_for_GPUs.html https://jax.readthedocs.io/en/latest/Custom_Operation_for_GP... I haven't worked with float4, but can imagine that new numerical types would require some special handling. But I assume that's the case for any ml environment. But really you probably mean fixed point 4bit integer types? Looks like that has had at least some work done in Jax: https://github.com/google/jax/issues/8566 https://github.com/google/jax/issues/8566
- jeffbee 3y agoThe H100 costs 10x on AWS compared to the TPUv5e on GCP. And it is apparently 5x faster in GPT training. Which makes the headline sort of backwards.
- latency-guy2 3y agoCost, availability, and speed. Salaries more often than not, costs a hell of a lot more than compute and energy bills cost. If you have a 10k cluster, you're probably a mega-corp running a few teams sharing the resources running experiments in parallel to each other and themselves. Now this is an very much so an overestimate on the actual cost of compute, the $10/hr/GPU cost is in all likelihood still significantly cheaper than the cash you pay to your researchers, and all the things that are needed to support them. Unless your team of researchers is about 10 people running very efficiently and making use of ALL the GPUs each hour of every day.
- d3w4s9 3y agoWhile power consumption is important, it is never a top priority when we talk about training these huge models. Raw performance matters more than anything else. And just because the article doesn't mention the aspect you care about does not make it a clickbait.
- jeffbee 3y agoWait, why? Is it more important to organizations to wait a little less time, or to spend less money?
- ss1996 3y agoThe cost of power (electricity) is much lesser than the cost of time.
- jeffbee 3y agoThat's not abundantly clear. Some of the facilities in the article cost $100k/hr.
- ss1996 3y agoThe cost of electricity for running a single RTX 4090 a year: $800 (calculation below). Cost to the company of the median developer: $200,000. Assume a 20 cent cost of 1 kWh, then a RTX 4090 with a 450W power draw for a year costs: 0.45 * 24 * 365 * 0.20 = $788.4
- ss1996 3y agoMoreover, considering the cost of the card itself at $1,600, the electricity expenses become relatively minor. In fact, for the price of the card, you could operate it non-stop for two years. This highlights how the cost of electricity is a small factor in the overall picture.
- 3y ago