8 ms·
Nvidia Hopper Sweeps AI Inference Benchmarks in MLPerf Debut
- IceHegel 4y agoI get why the 450% increase on BERT is the headline, but I think it's even more impressive that the smallest increase vs. A100 is still 50% (for ResNet). You don't get these types of generational improvements in the CPU world these days. In some ways though,it’s better that the gains go to the GPU because, after years of talk, we actually do seem to be in the middle of some really awesome ML developments.
- alchemist1e9 4y agoAs a favor could you list what you think are the most important ML developments recently? It’s hard to keep up and I have a sense from your comment you might have a condensed summary list that might be helpful.
- mitchellgoffpc 4y agoSome neat results from the last six months or so: - Significantly-improved diffusion models (DALL-E 2, Midjourney, Stable Diffusion, etc) - Diffusion models for video (see https://video-diffusion.github.io/ https://video-diffusion.github.io/, this paper is from April but I expect to see a lot more published research in this area soon) - OpenAI Minecraft w/VPT (first model with non-zero success rate at mining diamonds in <20min) - AlphaCode (from February, reasonably high success rate on solving competitive programming problems) - Improved realism and scale for NeRFs (see https://dellaert.github.io/NeRF22/ https://dellaert.github.io/NeRF22/ for some cool examples from this year’s CVPR) - Better sample efficiency for RL models (see https://arxiv.org/abs/2208.07860 https://arxiv.org/abs/2208.07860 for a recent real-world example)
- IceHegel 4y agoThis would basically be my list. I'd add GPT3 & Github Copilot, which my team and I use professionally. It's far from perfect, but it's a great GSD tool especially for stuff like Regex, bash scripts, and weird APIs.
- riku_iki 4y ago> - AlphaCode (from February, reasonably high success rate on solving competitive programming problems) reasonably high was 50% on 10 attempts, meaning success rate on first attempt can be as low as 5%, out of which who knows how many were leaked to training data.
- XorNot 4y agoThis is pretty big news when you look at the numbers floating around time to train AIs like Stable Diffusion - which is by far the most disruptive change that's happened recently. 150,000 hours just went down to 100,000 hours if we take the conservative estimate. Not fast enough (not at the price point) but getting there: of course since it's an entirely parallel problem, the real metric is cost-per-model since in reality you'd just buy time from the cloud. I figure we're going to start seeing some big changes once cost-per-model puts it within reach of the hobbyist. If I today could buy a custom model for say, $1000 - that's accessible to the hobbyist experimentalist.
- justatdotin 4y ago> by far the most disruptive change that has been disclosed :/
- colordrops 4y agoIt's disruptive because it has been disclosed.
- dannyw 4y agoEven $10k is enough for a hobbyist group or crowd funding.
- ipsum2 4y agoStable diffusion is on the order of millions of dollars.
- tehsauce 4y agoIt was about $200k in compute.
- moonshotideas 4y agoIs there a breakdown of the costs available to view?
- lunixbochs 4y agoRemember they also went from 400W on A100 to 700W on H100
- Cacti 4y agowow
- IceHegel 4y agoTrue, and a process shrink from TSMC 7nm to 4nm.
- pclmulqdq 4y agoDid they move to vertical power delivery for this generation? That would explain the increase in power delivery to the chip.
- sanxiyn 4y agoI don't think anyone has vertical power delivery yet.
- bambax 4y agoWhat is vertical power?
- pclmulqdq 4y agoVertical power delivery involves putting the voltage regulators on the opposite side of the circuit board of the chip, so that high-current power supplies travel a minimum distance. This reduces both resistance and inductance of the paths that power takes to the transistors, which means significantly less loss. That, in turn, means less heat related to those losses and less infrastructure to prevent those losses (fewer bypass capacitors, etc), so more of the thermal and area budgets can go to compute.
- bambax 4y agoAaah, opposite side as in... under side? Now I get it, thanks!!
- DavidSJ 4y agoNote that from their chart it’s 350% faster, i.e. 4.5x as fast (they say, as is unfortunately common, “4.5x faster”).
- KingOfCoders 4y ago(non native speaker) Instead of the "correct" 4.5x as fast?
- iforgotpassword 4y agoBoth things are correct in the grammatical sense, but they have different meaning, and in the Nvidia case, "4.5x as fast" would have been correct, or they should have used the less impressive sounding "3.5x faster". It might be easier to grasp if you use 1x. E.g. "it's 1x as fast" means it's exactly the same. You're applying the multiplier directly to get the new speed. 0.5x as fast means it's half as fast. "It's 1x faster" means you add the result of the multiplication to the initial value. So it's twice as fast. 0.5x faster still means it's faster, you add 50% of the initial speed. This way you cannot really express that something got slower, except by using negative values, which might be rather confusing. I think this might also work in most other languages originating in Europe.
- Nimitz14 4y agoI have never heard anybody use "x times faster" in the sense you mean here. To my ears, "3 times faster" has the same meaning as "3 times as fast".
- benplumley 4y agoHow about "50% faster" and "50% as fast"? If these are different (which to my ear they clearly are), then "200% faster" and "200% as fast" are different too. Naturally it's all quite ambiguous so whoever's writing the press release can pick the more impressive meaning in each case.
- probably_a_gpt 4y ago> we actually do seem to be in the middle of some really awesome ML developments You could say we are mid-journey…
- aliljet 4y agoI'm really curious what will happen with the export ban here. Inevitably, the blow back will be alternatives that emerge in markets that are not the United States. I'm curious if there are any non-US GPU contenders on the horizon?
- IceHegel 4y agoThe ban makes sense given ML compute is dual-use, but I wonder how hard it would be to train say, a missile targeting system, on AWS. I wonder if any nationstate would take the risk of using the public cloud for classified work as a way around sanctions.
- cjbgkagh 4y agoI don’t think it’s dual use for ML, I think it’s more about edge compute. Someone correct me if I’m wrong but I think radar resolution is compute bound so an order of magnitude or two could defeat stealth.
- beebmam 4y agoThe export ban applies to China only. Source: https://www.reuters.com/technology/nvidia-says-us-has-imposed-new-license-requirement-future-exports-china-2022-08-31/ https://www.reuters.com/technology/nvidia-says-us-has-impose...
- lajamerr 4y agoASICs/Google TPU like things for AI inference/training will be easy to replicate and will have 95% of the power of existing solutions or maybe even improvements. Because they are simple in design. As far as dedicated GPUs. There is so much additional work in the software side/driver side of thing unless there was a big concerted effort by a single company to build a solid foundation I see it as unlikely. Look at Intel struggling to get their GPUs out there. Dedicated accelerated cards just for ML though? I believe some already exist/more will come.
- noogle 4y agoAnd the biggest push to invest in re-inventing a software stack is a legal ban on access to the existing software. If previously building a separate CUDA-like framework competed against paying a bit more to Nvidia, now this decision is infinitely simpler, since using Nvidia products is just out of the question.
- modeless 4y agoTSMC 5/4nm strikes again. Everything fabricated on this process is great.
- KingOfCoders 4y agoIt shows how important it was to launch M1 on 5nm (great launch). Intel should have also waited for Arc to launch (poor launch) on 5nm.
- adgjlsfhk1 4y agoarc's problems seem to mostly be software.
- KingOfCoders 4y agoYes and no. Even with new DX versions it seems only a "meh" reentry to the GPU market - with older versions it looks like a catastrophe to me. The M1 was a "wow" entry.
- nisegami 4y agoIf my understanding is correct, "nm" numbers across fabs don't really compare very well. So 5nm working out well for TSMC doesn't say anything about 5nm for Intel.
- modeless 4y agoIntel Arc is fabricated at TSMC :-O
- bottlepalm 4y agoHats off to Nvidia for continually pushing things forward!
- nl 4y agoOut of interest I've been running a bunch of the huggingface version of StableDiffusion using the M1 accelerated branch on my M1 Max[1]. I'm getting 1.54 it/s compared to 2.0 it/s for a Nvidia T4 Tesla on Google Collab. T4 Tesla gets 21,691 queries/second for for ResNet, compared to 81,292 q/s for the new H100, 41,893 q/s for the A100 and 6164 q/s for the new Jetson. So you can expect maybe 15,000 q/s on a M1 Max. But some tests seem to indicate a lot less[2] - not sure what is happening there. [1] Setup like this: https://github.com/nlothian/m1_huggingface_diffusers_demo https://github.com/nlothian/m1_huggingface_diffusers_demo [2] https://tlkh.dev/benchmarking-the-apple-m1-max#heading-resnet50-inference https://tlkh.dev/benchmarking-the-apple-m1-max#heading-resne...
- deleted 4y ago[deleted]
- machinekob 4y agoCUDA vs MPS happen as apple software is always behind hardware by a lot.
- gianela 4y ago
- fxtentacle 4y agoI don't get why this is news now. The H100 diagrams have been online for weeks/months. I remember being confused by them when I researched A100 options in July.
- loser777 4y agoIf you can turn whitepaper diagrams and throughput claims into accurate performance estimates in your head, I suggest you embark on a lucrative career path as a world-leading computer architect immediately if you haven't already done so
- fxtentacle 4y agoI think I'd rather stay in my lowly AI research basement ;)
- deleted 4y ago[deleted]
- tmaly 4y agoIf these H100s were priced for retail/hobby I think we could see some very interesting innovations.