10 ms·
AMD Alveo V70 AI inference accelerator card
- westmeal 4y agoNo price is listed on their site so I'm assuming its gonna be stupid expensive, but if anyone knows would you mind posting?
- ChuckNorris89 4y agoProbably because it's a product aimed at datacenters and cloud providers who work directly with Xilinx/AMD to develop it, so they already know the price.
- dragontamer 4y agoPrevious AMD/Xilinx Alveo are in the $5000 to $20,000 range USD. I'd assume somewhere around there, or maybe even a bit higher. EDIT: 75W is a smaller card than I expected. "Inference" also usually means "cheaper". so maybe we can be optimistic with $5000-ish ?? Anyone shocked by the price, remember that this is an FPGA-line from Xilinx. Not a GPU from Radeon. Expect very high prices.
- dhruvdh 4y agoIt’s 1,995$ - I tried to order when it was announced.
- dhruvdh 4y agoI went through the checkout flow earlier, it was 1,995$ pre tax.
- mcilie 4y agoIt costs 1995 if you look at the "order now" section
- wyldfire 4y agoThe price is shown as $1,995.00 + tax&shipping for A-V70-P16G-ES3-G.
- novaRom 4y agoWhat TOPS means exactly in "... TOPS*|(INT8) 404 ..." ?
- capableweb 4y agoTOPS - Trillions of Operations Per Second, used as a benchmark to figure out the performance of the accelerator. In my experience, mostly a marketing number, higher TOPS doesn't actually mean it'll be faster than something with a lower TOPS. As always, you need to do your own benchmarks with your use case in mind.
- novaRom 4y agoWhat kind of operations is not clear. Wether it's a simple logic operation or a FMA is big difference.
- calaphos 4y agoI assume it's int8 operations (FMA would count as 2). At least that's the case with FLOPs, TOPs for AI accelerators is basically the same measure, with the changed acronym reflecting the non float data format.
- wyldfire 4y agoAMD XDNA – Versal AI Core / 2nd-gen AIE-ML tiles Are these programmable by the end-user? The "software programmability" section describes "Vitis AI" frameworks supported. But can we write our own software on these? Is this card FPGA-based? EDIT: [1] more info on the AI-engine tiles: scalar cores + "adaptable hardware (FPGA?)" + {AI+DSP}. [1] https://www.xilinx.com/products/technology/ai-engine.html https://www.xilinx.com/products/technology/ai-engine.html
- djmips 4y agoSays RDNA based which is AMD's GPU tech.
- derefr 4y agoIt's very likely FPGA-based; Xilinx is an FPGA company. This is being pitched as an "AI accelerator", but "Alveo" as a product line existed before AMD's acquisition of Xilinx, and other "Alveo" products exist (https://www.xilinx.com/products/boards-and-kits/alveo.html https://www.xilinx.com/products/boards-and-kits/alveo.html) that are marketed for other purposes, while really just being Xilinx FPGAs pre-programmed to perform specific other tasks, with some domain-specific DSPs + interconnects around the edges. It's possible that AMD could have reworked an existing Xilinx design to incorporate RDNA chiplets in place of some of the FPGA-gate-grid chiplets, creating a heterogeneous mesh; but I find it just as likely that AMD just took their VLSI for an RDNA core and loaded it onto the existing FPGA.
- typon 4y agoIt's not a traditional FPGA chip (lots of luts and flip flops). The "AI Engine" is basically hardened chiplets that are working alongside soft logic chiplets and I/O. This is how they're able to get their performance/power numbers
- messe 4y ago> High-Density Video Decoder**: 96 channels of 1920x1080p > [...] > **: @10 fps, H.264/H.265 Is 10 fps a standard measure for this kind of thing?
- novaRom 4y ago10 fps should be fast enough to provide input tensors for real time inference with small scale transformers / convolutional nets.
- scottlamb 4y agoMaybe. Running inference at 10 fps is probably plenty. But that doesn't mean you only have to do 10 fps of H.264/H.265 decoding. I think the most common scenario is for the input video to be e.g. 30 fps with mostly P frames that each depend on the prior frame in a chain. In that case, you need to decode almost [1] 30 fps to get 10 fps of evenly spaced frames to process. [1] You could skip the last P frame before an IDR frame, but that doesn't buy you much.
- zamadatix 4y agoIf your source is 96 YouTube videos sure, if it's 96 CCTV cameras it's different.
- scottlamb 4y agoStill depends. As it happens, I'm developing my own open source NVR software, [1] so I know a bit about this. Some cameras are fairly good about this, supporting the following features: * "Temporal SVC", in which the frame dependencies are structured so you can discard down to 1/2 or 1/4th of the nominal frame rate and still decode the remainder. * Three output streams, which you could configure for say forensics (high-bandwidth/high-resolution/high-fps), inference (mid-bandwidth/mid-resolution/low-fps), and viewing multiple streams / over mobile networks (low-bandwidth/low-resolution/mid-fps). * On-camera ML tasks too. (Although I haven't seen one that lets you upload your own model.) But other cameras are less good. E.g. some Reolinks [2] only support two streams, and the "sub" stream is fixed at 640x352, which is uncomfortably low. Your inference network may not take more resolution than that, but even if not, you might want to crop down to the area of interest (where there's motion and/or where the user has configured an alert) to improve quality. (You probably wouldn't pair that cheap Reolink camera with this expensive inference card, but the point stands in general.) Even the "better" cameras' timestamp handling is awful, so it's hard to reliably match up the main stream, sub stream, analytics output, and wall clock time. Given that limitation it'd be desirable to just use the main stream for everything but the on-NVR transcoding's likely unaffordable. [1] https://github.com/scottlamb/moonfire-nvr https://github.com/scottlamb/moonfire-nvr [2] https://github.com/scottlamb/moonfire-nvr/wiki/Cameras:-Reolink#reolink-rlc-410-hardware-version-ipc_3816m https://github.com/scottlamb/moonfire-nvr/wiki/Cameras:-Reol...
- h2odragon 4y agoDouglas Adams said we'd have robots to watch TV for us. That seems to be the designed use case for this. 16gb RAM / 96 video channels ... I haven't done any of that work but it feels like they expect that "96" not to be fully used in practice.
- inetsee 4y agoI have no problem imagining a security camera application needing to monitor quite a few video channels.
- andy_ppp 4y ago/camera/state/g
- h2odragon 4y agoCertainly. I'm suspecting that doing much of anything with all 96 channels would really need more RAM, for most users.
- scottlamb 4y agoOn the inference accelerator? IIUC, the RAM is just to hold the model and whatever state it needs during a particular inference operation. I'm not an expert on ML but AFAIK 16 GiB is plenty. I suppose it'd also need to hold onto reference frames for the video decoding, but at 1080p with e.g. YUV420 (12 bits per pixel), you can hold a lot of those in 16 GiB. edit: e.g., 4 references for each of the 96 streams would take ~1 GiB. Even on the host, 16 GiB is fine for say an NVR. They don't need to keep a lot of state in RAM (or for that matter to do a lot of on-CPU computation either). I can run an 8-stream NVR on a Raspberry Pi 2 without on-NVR analytics. That's about its limit because the network and disk are on the same USB2 bus, but there's spare CPU and RAM.
- phkahler 4y ago>> I have no problem imagining a security camera application needing to monitor quite a few video channels. As a joke I sometimes tell people the automatic flushing toilets in public bathrooms work by having a little camera monitored by someone in a 3rd world country who remotely flushes as needed, while monitoring a whole lot of video feeds. They usually don't buy it, but will often acknowledge that our world is uncomfortably close to having stuff like become reality.
- hyuen 4y agoI won't even take a look at the numbers unless they show a PyTorch model running on it, the problem is the big disconnect between HW and SW, realistically, have you ever seen any off-the shelf model running on something other than NVidia?
- deleted 4y ago[deleted]
- Narew 4y agoIt's for inference only not training. In this use case, there is lots of device that's not Nvidia. For server you have Google tpu, for more close to public there is the Apple Neural Engine for example.
- kombine 4y agoThat. I work on the research side and I am still waiting for non-NVIDIA hardware for training deep models.
- frozenport 4y agoGoogle's TPU comes to mind
- hhh 4y agoI've ran models on Apple HW, Raspberry Pi, random CPUs, NVIDIA GPUs, TPUs. Waiting to get my hands on Tenstorrent gear.
- kombine 4y agoThis is inference only. AMD should invest into the full AI stack starting from training. For this they need a product comparable to NVIDIA 4090, so that entry level researchers could use their hardware. Honestly, I don't know why AMD aren't doing that already, they are best positioned to do that in the industry landscape.
- Hooray_Darakian 4y ago> AMD should invest into the full AI stack starting from training. https://www.amd.com/en/graphics/servers-solutions-rocm-ml https://www.amd.com/en/graphics/servers-solutions-rocm-ml > For this they need a product comparable to NVIDIA 4090, so that entry level researchers could use their hardware. Why is a high end product a requirement for entry level research?
- jzwinck 4y agoBecause high end research uses a fleet of them, not just one.
- kombine 4y ago4090 (or 3090, 1080Ti and so on) is a high-end consumer GPU, but at the same time it is an entry level GPU for AI researchers. Don't forget that workstation cards (RTX 8000) let alone server-grade GPUs such as A100 are an order of magnitude more expensive.
- gymbeaux 4y agoI was doing ML stuff on a GTX 1060 a few years ago. As with everything, it depends on what you’re doing.
- p1esk 4y agoChances of you publishing something in ML improve proportionally to the amount of hardware you have access to. Or to put it another way, the less hardware you have, the smarter you have to be to publish something in ML.
- p1esk 4y agoHow much memory does it have?
- kmeisthax 4y agoEvery time I hear about an AI accelerator, I get really excited, then it turns out to be inference only.
- psychphysic 4y agoDid I miss something did AMD buy Xilinx? Makes sense I suppose after Intel bought Altera. Who owns lattice?
- AnonMO 4y agoIn the time it took you to write this you could have searched your two question on google clicked the first wiki link and got your answer
- psychphysic 4y agoBah! I'll go back to voicing all my thoughts to chagpt
- tgtweak 4y agoSad to see amds ROCm efforts essentially abandoned. They were close to universal interop for cudnn and cuda on amd (and other!) Architectures. Hopefully Intel takes a stab at it with their ARC line out now.
- Roark66 4y agoIf this is based on fpga tech (xilinx) I don't think it will have a cost/benefit edge over asics. Why not do their own TPU like Google did? Nowadays even embedded MCUs come with AI accelerators (last I heard was 5TOPS in a banana pi-cm4 board - that is sufficient for object detection stuff and perhaps even more).